awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
breezedeus avatar

breezedeus/Pix2Text

0
View on GitHub↗
3,012 نجوم·261 تفرعات·Jupyter Notebook·mit·19 مشاهداتp2t.breezedeus.com↗

Pix2Text

Pix2Text is an optical character recognition system and document conversion tool designed to transform images and PDFs into Markdown. It functions as a multilingual OCR engine supporting over 80 languages, a LaTeX formula recognizer for mathematical notations, and a parser integrated with vision language models.

The project utilizes a hybrid pipeline to separate plain text from mathematical formulas and tabular structures within a single pass. It converts recognized formulas into LaTeX expressions and transforms detected tables and layouts into structured Markdown formatting.

The system includes a command line interface for document conversion and a local HTTP web API for programmatic image processing. It supports GPU acceleration to increase model inference speed.

Features

  • Document to Markdown Converters - Transforms images and PDFs into formatted Markdown documents by extracting layout, tables, and formulas.
  • Image-to-LaTeX Converters - Automatically transcribes visual mathematical notation from images into structured LaTeX code.
  • OCR Pipelines - Implements a hybrid OCR pipeline that separates plain text and mathematical formulas in a single pass.
  • Multilingual Glyph Mappings - Uses extended language packs to map visual glyphs from over 80 different languages to digital text.
  • Multilingual OCR Systems - Implements an OCR system capable of processing diverse character sets for over 80 global languages.
  • Document Layout Analysis - Analyzes page layouts to distinguish between text blocks, tables, and mathematical formulas before recognition.
  • Optical Character Recognition - Extracts printed text from images across over 80 different global languages into digital format.
  • Multilingual Text Recognition - Recognizes and converts printed characters from over 80 different languages into digital text.
  • Optical Character Recognitions - Uses optical character recognition to convert images containing plain text into digital strings.
  • Formula Extractors - Isolates mathematical notation from document images for conversion into digital LaTeX expressions.
  • Formula Recognition Engines - Translates images containing mathematical formulas into standardized LaTeX code.
  • OCR Document Conversion - Provides a command line interface to convert images and PDFs into structured Markdown via OCR.
  • GPU-Accelerated Inference - Utilizes GPU hardware acceleration to increase the inference speed of image and document processing models.
  • Model Serving APIs - Wraps the recognition engine in a local web server to provide OCR capabilities via a REST API.
  • Image-to-Markdown Table Generators - Recognizes tabular data within images and converts it into formatted Markdown tables.
  • Vision-Based Document Parsers - Uses multimodal vision language models to interpret document layouts and structural organization.
  • PDF to Markdown Conversion - Transforms PDF documents into structured Markdown files while preserving original text and table structures.
  • Document Reconstruction Serializers - Transforms recognized layout, tables, and formulas into a structured Markdown format for document reconstruction.
  • OCR Integration APIs - Provides a local HTTP service and API for programmatically processing images to extract text and formulas.
  • Tabular Data Extraction - Detects tables within documents and extracts content while preserving the original tabular structure.
  • OCR Web Services - Provides a local HTTP web API for programmatic image processing and OCR result retrieval.
  • Service Hosting - Exposes the OCR engine as a local HTTP service for programmatic image processing via API.

سجل النجوم

مخطط تاريخ النجوم لـ breezedeus/pix2textمخطط تاريخ النجوم لـ breezedeus/pix2text

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ Pix2Text

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Pix2Text.
  • opendatalab/pdf-extract-kitالصورة الرمزية لـ opendatalab

    opendatalab/PDF-Extract-Kit

    9,724عرض على GitHub↗

    PDF-Extract-Kit is a document extraction toolkit designed to convert PDF documents into structured formats such as Markdown, HTML, and LaTeX. It functions as a multi-stage parsing framework that combines a document layout analyzer, a formula recognition engine, an OCR text extractor, and a table extraction system. The project focuses on recovering complex document elements by translating images of mathematical formulas and tabular structures into editable source code. It utilizes model-driven layout analysis to identify structural elements in reports and textbooks while ignoring noise like wa

    Python
    عرض على GitHub↗9,724
  • opendataloader-project/opendataloader-pdfالصورة الرمزية لـ opendataloader-project

    opendataloader-project/opendataloader-pdf

    25,769عرض على GitHub↗

    This project is a PDF data extraction tool and document preprocessor designed to convert PDF files into structured formats such as Markdown, JSON, and HTML. It functions as an OCR document parser for scanned files, an accessibility automator for generating PDF/UA compliant metadata, and a loader for AI orchestration frameworks like LangChain. The software distinguishes itself through specialized handling of complex document elements, including the conversion of mathematical formulas into LaTeX and the generation of natural-language descriptions for charts and images. It utilizes recursive seg

    Javaa11yaccessibilityai
    عرض على GitHub↗25,769
  • lukas-blecher/latex-ocrالصورة الرمزية لـ lukas-blecher

    lukas-blecher/LaTeX-OCR

    16,190عرض على GitHub↗

    LaTeX-OCR is a specialized optical character recognition system designed to identify and transcribe complex mathematical symbols and their spatial relationships from images. It functions as a machine learning engine that converts visual representations of equations into structured LaTeX code for use in technical documentation and academic typesetting. The project utilizes a hierarchical vision-based encoding and autoregressive sequence decoding architecture to process input images and generate mathematical notation token by token. Beyond its core recognition capabilities, the system provides

    Pythondatasetdeep-learningim2latex
    عرض على GitHub↗16,190
  • rednote-hilab/dots.ocrالصورة الرمزية لـ rednote-hilab

    rednote-hilab/dots.ocr

    7,695عرض على GitHub↗

    dots.ocr is a suite of software utilities for document layout analysis, multilingual optical character recognition, and scene text digitization. It functions as an engine for extracting digital text and structured layout data from images and PDFs across various human scripts. The project includes a specialized transformer for converting charts, diagrams, and chemical formulas from raster images into scalable vector graphics. It also provides a pipeline to transform extracted text and structural layout from documents and web screenshots into formatted Markdown files. The system covers capabil

    Python
    عرض على GitHub↗7,695
عرض جميع البدائل الـ 30 لـ Pix2Text→

الأسئلة الشائعة

ما هي وظيفة breezedeus/pix2text؟

Pix2Text is an optical character recognition system and document conversion tool designed to transform images and PDFs into Markdown. It functions as a multilingual OCR engine supporting over 80 languages, a LaTeX formula recognizer for mathematical notations, and a parser integrated with vision language models.

ما هي الميزات الرئيسية لـ breezedeus/pix2text؟

الميزات الرئيسية لـ breezedeus/pix2text هي: Document to Markdown Converters, Image-to-LaTeX Converters, OCR Pipelines, Multilingual Glyph Mappings, Multilingual OCR Systems, Document Layout Analysis, Optical Character Recognition, Multilingual Text Recognition.

ما هي البدائل مفتوحة المصدر لـ breezedeus/pix2text؟

تشمل البدائل مفتوحة المصدر لـ breezedeus/pix2text: opendatalab/pdf-extract-kit — PDF-Extract-Kit is a document extraction toolkit designed to convert PDF documents into structured formats such as… opendataloader-project/opendataloader-pdf — This project is a PDF data extraction tool and document preprocessor designed to convert PDF files into structured… lukas-blecher/latex-ocr — LaTeX-OCR is a specialized optical character recognition system designed to identify and transcribe complex… rednote-hilab/dots.ocr — dots.ocr is a suite of software utilities for document layout analysis, multilingual optical character recognition,… run-llama/liteparse — A fast, helpful, and open-source document parser. tesseract-ocr/tessdata — This repository provides the pre-trained neural network and legacy data files used by Tesseract to recognize and…