awesome-repositories.comKategorienBlog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
C

CatchTheTornado/pdf-extract-api

0
View on GitHub↗
0 Stars·0 Forks·6 Aufrufe

Pdf Extract Api

Features

  • Data Processing - API for document extraction using modern OCR and local models.
  • Data Processing Tools - API for document extraction using modern OCR and local models.

Star-Verlauf

Star-Verlauf für catchthetornado/pdf-extract-apiStar-Verlauf für catchthetornado/pdf-extract-api

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Open-Source-Alternativen zu Pdf Extract Api

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Pdf Extract Api.
  • jpmens/joAvatar von jpmens

    jpmens/jo

    4,868Auf GitHub ansehen↗

    Jo is a command-line utility designed to construct and manipulate JSON objects and arrays directly from shell arguments and standard input. It functions as a data processing tool that transforms raw input into structured formats, enabling the generation of complex payloads for APIs, configuration files, and automated data pipelines. The tool distinguishes itself through its ability to resolve hierarchical data structures using delimiter-based path definitions and its integrated type-inference engine, which automatically casts input values into native boolean, numeric, or null types. Users can

    C
    Auf GitHub ansehen↗4,868
  • argilla-io/distilabelAvatar von argilla-io

    argilla-io/distilabel

    3,277Auf GitHub ansehen↗

    Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

    Python
    Auf GitHub ansehen↗3,277
  • 599yongyang/datasetloom5

    599yongyang/DatasetLoom

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • allenai/olmocrAvatar von allenai

    allenai/olmocr

    17,396Auf GitHub ansehen↗

    Olmocr is a distributed document processing framework designed to convert PDF and image files into structured markdown. It functions as a vision-based document parser that utilizes multimodal neural networks to interpret complex visual layouts and translate them into standardized text representations. The system operates as a remote inference orchestrator, offloading heavy document analysis tasks to external servers or cloud APIs to minimize local computational requirements. By employing a stateless worker architecture, it decouples document ingestion from inference, allowing for the distribu

    Python
    Auf GitHub ansehen↗17,396
Alle 30 Alternativen zu Pdf Extract Api anzeigen→

Häufig gestellte Fragen

Was sind die Hauptfunktionen von catchthetornado/pdf-extract-api?

Die Hauptfunktionen von catchthetornado/pdf-extract-api sind: Data Processing, Data Processing Tools.

Welche Open-Source-Alternativen gibt es zu catchthetornado/pdf-extract-api?

Open-Source-Alternativen zu catchthetornado/pdf-extract-api sind unter anderem: jpmens/jo — Jo is a command-line utility designed to construct and manipulate JSON objects and arrays directly from shell… chatdoc-com/ocrflux — OCRFlux is a lightweight yet powerful multimodal toolkit that significantly advances PDF-to-Markdown conversion,… 599yongyang/datasetloom. allenai/olmocr — Olmocr is a distributed document processing framework designed to convert PDF and image files into structured… bytedance/dolphin — Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital… argilla-io/distilabel — Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable…