awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
C

ckorzen/pdf-text-extraction-benchmark

0
View on GitHub↗
0 estrellas·0 forks·3 vistas

Pdf Text Extraction Benchmark

This project is about benchmarking and evaluating existing PDF extraction tools on their semantic abilities to extract the body texts from PDF documents, especially from scientific articles. It provides (1) a benchmark generator, (2) a ready-to-use benchmark and (3) an extensive evaluation, with…

Features

  • PDF Processing Tools - Benchmark for evaluating PDF text extraction tools.

Historial de estrellas

Gráfico del historial de estrellas de ckorzen/pdf-text-extraction-benchmarkGráfico del historial de estrellas de ckorzen/pdf-text-extraction-benchmark

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a Pdf Text Extraction Benchmark

Proyectos open-source similares, clasificados según cuántas características comparten con Pdf Text Extraction Benchmark.
  • frooodle/stirling-pdfAvatar de Frooodle

    Frooodle/Stirling-PDF

    81,168Ver en GitHub↗

    Stirling-PDF is a web-based PDF management suite used for editing, merging, splitting, and converting PDF documents. It functions as a self-hosted document manager, providing a centralized interface for users to manipulate files on a private server. The system features a workflow automation engine that allows for the creation of processing pipelines to handle large volumes of documents without writing custom code. It also includes an optical character recognition tool to convert scanned PDFs into searchable and editable text. Access is managed through single sign-on integration and OIDC comp

    Java
    Ver en GitHub↗81,168
  • guaguastandup/zotero-pdf2zhAvatar de guaguastandup

    guaguastandup/zotero-pdf2zh

    3,008Ver en GitHub↗

    zotero-pdf2zh is a translation system and Zotero plugin designed to convert academic PDF papers into target languages using large language model APIs and machine translation services. It functions as an LLM document translator that integrates directly into Zotero libraries to process research documents. The tool distinguishes itself by generating bilingual PDFs with parallel layouts for side-by-side comparison of original and translated text. It also includes a mobile layout optimizer that crops double-column PDFs and stacks them vertically to improve readability on narrow screens. The syste

    Python
    Ver en GitHub↗3,008
  • applicaai/digital-born-pdf-scannerA

    applicaai/digital-born-pdf-scanner

    0Ver en GitHub↗

    Many of PDF files that we have downloaded are digital-born, that is contain easily accessible text layer that PDF viewers use to display text. Some are definitely scanned documents, that do not have any text layer at all, some are searchable OCR-processed scans that contain a lot of hidden text.

    Ver en GitHub↗0
  • deepdoctection/deepdoctectionAvatar de deepdoctection

    deepdoctection/deepdoctection

    3,181Ver en GitHub↗

    A Repo For Document AI

    Python
    Ver en GitHub↗3,181
Ver las 16 alternativas a Pdf Text Extraction Benchmark→

Preguntas frecuentes

¿Qué hace ckorzen/pdf-text-extraction-benchmark?

This project is about benchmarking and evaluating existing PDF extraction tools on their semantic abilities to extract the body texts from PDF documents, especially from scientific articles. It provides (1) a benchmark generator, (2) a ready-to-use benchmark and (3) an extensive evaluation, with…

¿Cuáles son las características principales de ckorzen/pdf-text-extraction-benchmark?

Las características principales de ckorzen/pdf-text-extraction-benchmark son: PDF Processing Tools.

¿Qué alternativas de código abierto existen para ckorzen/pdf-text-extraction-benchmark?

Las alternativas de código abierto para ckorzen/pdf-text-extraction-benchmark incluyen: frooodle/stirling-pdf — Stirling-PDF is a web-based PDF management suite used for editing, merging, splitting, and converting PDF documents.… guaguastandup/zotero-pdf2zh — zotero-pdf2zh is a translation system and Zotero plugin designed to convert academic PDF papers into target languages… applicaai/digital-born-pdf-scanner — Many of PDF files that we have downloaded are digital-born, that is contain easily accessible text layer that PDF… deepdoctection/deepdoctection — A Repo For Document AI. apache/pdfbox — Licensed to the Apache Software Foundation (ASF) under one or more contributor license agreements. See the NOTICE file… allenai/pawls — Demo Server | Video Tutorial | Paper.