awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
getomni-ai avatar

getomni-ai/zerox

0
View on GitHub↗
12,241 نجوم·846 تفرعات·TypeScript·MIT·5 مشاهداتgetomni.ai/ocr-demo↗

Zerox

Zerox is a multimodal document parser and OCR tool that uses vision models to convert PDF files and images into structured Markdown text. It functions as a visual layout extraction engine, leveraging large multimodal models to digitize documents while maintaining their original structural formatting.

The system differentiates itself through the use of coordinate-based element mapping and multimodal layout analysis to identify structural elements like tables, charts, and headers. It utilizes rasterization to convert vector PDF pages into high-resolution bitmaps, ensuring consistent input for the vision models used to synthesize the final Markdown output.

The tool covers a broad range of document digitization capabilities, including complex layout extraction and vision-based OCR. It processes visual document representations to interpret the spatial relationship between text and data, converting them into machine-readable formats.

Features

  • PDF to Markdown Converters - Transforms PDF files into Markdown text using vision models to preserve the original layouts, tables, and charts.
  • Multimodal Vision Models - Identifies structural elements like tables and headers by processing document images through a large multimodal model.
  • Automated Digitization Engines - Converts physical or digital document scans into machine-readable formats using multimodal models to identify structural elements.
  • Multimodal Layout Analysis - Uses large vision models to identify structural document elements like tables and headers from image data.
  • Structured Document Extraction - Converts complex PDF files into structured Markdown while preserving tables, charts, and the original page formatting.
  • AI Vision OCR Tools - Provides a document extraction system that uses vision models to convert PDF files and images into structured Markdown text.
  • Vision-Based Document Parsers - Provides a parser that uses multimodal vision models to interpret document layouts and convert them into structured text.
  • Visual-to-Markdown Pipelines - Transforms visual document representations into structured text by mapping identified coordinates to Markdown formatting syntax.
  • PDF to Markdown Conversion - Transforms visual document representations into structured Markdown text by mapping spatial coordinates to formatting syntax.
  • Multimodal Document Ingestion - Uses vision-capable large language models to interpret and convert visual document representations into clean, structured text.
  • Layout-Aware Extraction - Parses documents with intricate formatting to maintain the spatial relationship between text, headers, and tabular data.
  • Optical Character Recognition - Extracts raw character data from images using optical character recognition to supplement semantic structural formatting.
  • Element Discrimination Prompts - Uses specialized visual prompts to help models distinguish between standard body text and complex tabular data.
  • Multimodal Prompting - Employs specialized visual prompts to help the model distinguish between body text and complex tabular data.
  • Text Extraction and OCR - Recovers raw character data from images using OCR before applying semantic structural formatting.
  • Coordinate-Based Layout Mapping - Implements a coordinate-based mapping system to preserve the original document layout during the conversion process.
  • Document Layout Extraction - Identifies structural elements in documents through coordinate-based mapping and vision-model analysis.
  • Vector Rasterizers - Converts vector PDF pages into high-resolution bitmaps to provide consistent input for vision-based multimodal models.
  • PDF to Image Rendering - Converts vector-based PDF pages into high-resolution bitmaps to ensure compatibility with vision-based model inputs.
  • Data Processing - Zero-shot PDF OCR using vision-capable language models.
  • Data Processing Tools - Zero-shot PDF OCR utilizing vision-language models.

سجل النجوم

مخطط تاريخ النجوم لـ getomni-ai/zeroxمخطط تاريخ النجوم لـ getomni-ai/zerox

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ Zerox

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Zerox.
  • pymupdf/pymupdfالصورة الرمزية لـ pymupdf

    pymupdf/PyMuPDF

    9,086عرض على GitHub↗

    PyMuPDF is a comprehensive PDF manipulation library and document analysis tool. It serves as a text extraction tool, OCR engine, and image converter, providing a programmatic interface to edit, merge, split, and optimize PDF and Office documents. The project distinguishes itself through high-performance capabilities, including the use of C-bindings for low-level manipulation and parallelized page processing to accelerate workloads. It provides specialized conversion paths, such as transforming PDF content into Markdown for retrieval-augmented generation and large language model pipelines. It

    Pythondata-scienceepubextract-data
    عرض على GitHub↗9,086
  • opendatalab/pdf-extract-kitالصورة الرمزية لـ opendatalab

    opendatalab/PDF-Extract-Kit

    9,724عرض على GitHub↗

    PDF-Extract-Kit is a document extraction toolkit designed to convert PDF documents into structured formats such as Markdown, HTML, and LaTeX. It functions as a multi-stage parsing framework that combines a document layout analyzer, a formula recognition engine, an OCR text extractor, and a table extraction system. The project focuses on recovering complex document elements by translating images of mathematical formulas and tabular structures into editable source code. It utilizes model-driven layout analysis to identify structural elements in reports and textbooks while ignoring noise like wa

    Python
    عرض على GitHub↗9,724
  • run-llama/liteparseالصورة الرمزية لـ run-llama

    run-llama/liteparse

    10,782عرض على GitHub↗

    A fast, helpful, and open-source document parser

    Rustdocument-ocrdocument-processingocr
    عرض على GitHub↗10,782
  • bytedance/dolphinالصورة الرمزية لـ bytedance

    bytedance/Dolphin

    8,820عرض على GitHub↗

    Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital document images into machine-readable structured data. It functions as an LLM document parser, utilizing vision-language models to simultaneously predict spatial layout and text content. The system is designed as a concurrent document processor, employing parallel document parsing to process multiple elements across distributed compute nodes. This high-throughput approach reduces the total time required to convert large volumes of images into structured formats. The project covers

    Pythondocument-analysislayout-analysisocr
    عرض على GitHub↗8,820
عرض جميع البدائل الـ 30 لـ Zerox→

الأسئلة الشائعة

ما هي وظيفة getomni-ai/zerox؟

Zerox is a multimodal document parser and OCR tool that uses vision models to convert PDF files and images into structured Markdown text. It functions as a visual layout extraction engine, leveraging large multimodal models to digitize documents while maintaining their original structural formatting.

ما هي الميزات الرئيسية لـ getomni-ai/zerox؟

الميزات الرئيسية لـ getomni-ai/zerox هي: PDF to Markdown Converters, Multimodal Vision Models, Automated Digitization Engines, Multimodal Layout Analysis, Structured Document Extraction, AI Vision OCR Tools, Vision-Based Document Parsers, Visual-to-Markdown Pipelines.

ما هي البدائل مفتوحة المصدر لـ getomni-ai/zerox؟

تشمل البدائل مفتوحة المصدر لـ getomni-ai/zerox: pymupdf/pymupdf — PyMuPDF is a comprehensive PDF manipulation library and document analysis tool. It serves as a text extraction tool,… opendatalab/pdf-extract-kit — PDF-Extract-Kit is a document extraction toolkit designed to convert PDF documents into structured formats such as… run-llama/liteparse — A fast, helpful, and open-source document parser. bytedance/dolphin — Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital… pdf2htmlex/pdf2htmlex — pdf2htmlEX is a PDF to HTML converter that transforms documents into web pages while preserving the original layout,… facebookresearch/nougat — Nougat is a neural OCR system and LLM document parser designed to convert images of academic PDF documents into…