1 个仓库
AI-driven tools for cleaning and correcting text, notation, and structural elements within documents.
Distinct from LLM-Based Analysis: Distinct from LLM-Based Analysis: focuses on cleaning and merging content rather than analyzing code changes.
Explore 1 awesome GitHub repository matching devops & infrastructure · Document Content Refiners. Refine with filters or upvote what's useful.
Marker is an LLM-powered document parser and OCR pipeline designed to convert PDFs and unstructured files into structured markdown, JSON, and HTML. It functions as a data preprocessor that transforms complex documents into machine-readable formats while preserving tables, equations, and layout structures. The system utilizes large language models to refine OCR accuracy, clean mathematical notation, and merge fragmented tables across multiple pages. It employs model-based layout analysis to predict block types and bounding boxes, ensuring a more precise conversion of document elements. Capabi
Uses large language models to refine OCR accuracy, clean mathematical notation, and merge fragmented tables.