awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetÀ proposNotre méthodologiePresseServeur MCP
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
chatdoc-com avatar

chatdoc-com/OCRFlux

0
View on GitHub↗
2,514 stars·151 forks·Python·Apache-2.0·3 vues

OCRFlux

OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex layout handling, complicated table parsing and cross-page content merging.

Features

  • Data Processing - Multimodal toolkit for advanced PDF-to-Markdown conversion.
  • Data Processing Tools - Multimodal toolkit for advanced PDF-to-Markdown conversion.

Historique des stars

Graphique de l'historique des stars pour chatdoc-com/ocrfluxGraphique de l'historique des stars pour chatdoc-com/ocrflux

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Alternatives open source à OCRFlux

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec OCRFlux.
  • jpmens/joAvatar de jpmens

    jpmens/jo

    4,868Voir sur GitHub↗

    Jo is a command-line utility designed to construct and manipulate JSON objects and arrays directly from shell arguments and standard input. It functions as a data processing tool that transforms raw input into structured formats, enabling the generation of complex payloads for APIs, configuration files, and automated data pipelines. The tool distinguishes itself through its ability to resolve hierarchical data structures using delimiter-based path definitions and its integrated type-inference engine, which automatically casts input values into native boolean, numeric, or null types. Users can

    C
    Voir sur GitHub↗4,868
  • argilla-io/distilabelAvatar de argilla-io

    argilla-io/distilabel

    3,277Voir sur GitHub↗

    Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

    Python
    Voir sur GitHub↗3,277
  • 599yongyang/datasetloom5

    599yongyang/DatasetLoom

    0Voir sur GitHub↗
    Voir sur GitHub↗0
  • allenai/olmocrAvatar de allenai

    allenai/olmocr

    17,396Voir sur GitHub↗

    Olmocr is a distributed document processing framework designed to convert PDF and image files into structured markdown. It functions as a vision-based document parser that utilizes multimodal neural networks to interpret complex visual layouts and translate them into standardized text representations. The system operates as a remote inference orchestrator, offloading heavy document analysis tasks to external servers or cloud APIs to minimize local computational requirements. By employing a stateless worker architecture, it decouples document ingestion from inference, allowing for the distribu

    Python
    Voir sur GitHub↗17,396
Voir les 30 alternatives à OCRFlux→

Questions fréquentes

Que fait chatdoc-com/ocrflux ?

OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion, excelling in complex layout handling, complicated table parsing and cross-page content merging.

Quelles sont les fonctionnalités principales de chatdoc-com/ocrflux ?

Les fonctionnalités principales de chatdoc-com/ocrflux sont : Data Processing, Data Processing Tools.

Quelles sont les alternatives open-source à chatdoc-com/ocrflux ?

Les alternatives open-source à chatdoc-com/ocrflux incluent : jpmens/jo — Jo is a command-line utility designed to construct and manipulate JSON objects and arrays directly from shell… catchthetornado/pdf-extract-api. 599yongyang/datasetloom. allenai/olmocr — Olmocr is a distributed document processing framework designed to convert PDF and image files into structured… bytedance/dolphin — Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital… argilla-io/distilabel — Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable…