awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
C

ckorzen/pdf-text-extraction-benchmark

0
View on GitHub↗
0 stars·0 forks·7 views

Pdf Text Extraction Benchmark

This project is about benchmarking and evaluating existing PDF extraction tools on their semantic abilities to extract the body texts from PDF documents, especially from scientific articles. It provides (1) a benchmark generator, (2) a ready-to-use benchmark and (3) an extensive evaluation, with…

Features

  • PDF Processing Tools - Benchmark for evaluating PDF text extraction tools.

Star history

Star history chart for ckorzen/pdf-text-extraction-benchmarkStar history chart for ckorzen/pdf-text-extraction-benchmark

How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to Pdf Text Extraction Benchmark

Similar open-source projects, ranked by how many features they share with Pdf Text Extraction Benchmark.
  • frooodle/stirling-pdfFrooodle avatar

    Frooodle/Stirling-PDF

    81,168View on GitHub↗

    Stirling-PDF is a web-based PDF management suite used for editing, merging, splitting, and converting PDF documents. It functions as a self-hosted document manager, providing a centralized interface for users to manipulate files on a private server. The system features a workflow automation engine that allows for the creation of processing pipelines to handle large volumes of documents without writing custom code. It also includes an optical character recognition tool to convert scanned PDFs into searchable and editable text. Access is managed through single sign-on integration and OIDC comp

    Java
    View on GitHub↗81,168
  • guaguastandup/zotero-pdf2zhguaguastandup avatar

    guaguastandup/zotero-pdf2zh

    3,008View on GitHub↗

    zotero-pdf2zh is a translation system and Zotero plugin designed to convert academic PDF papers into target languages using large language model APIs and machine translation services. It functions as an LLM document translator that integrates directly into Zotero libraries to process research documents. The tool distinguishes itself by generating bilingual PDFs with parallel layouts for side-by-side comparison of original and translated text. It also includes a mobile layout optimizer that crops double-column PDFs and stacks them vertically to improve readability on narrow screens. The syste

    Python
    View on GitHub↗3,008
  • applicaai/digital-born-pdf-scannerA

    applicaai/digital-born-pdf-scanner

    0View on GitHub↗

    Many of PDF files that we have downloaded are digital-born, that is contain easily accessible text layer that PDF viewers use to display text. Some are definitely scanned documents, that do not have any text layer at all, some are searchable OCR-processed scans that contain a lot of hidden text.

    View on GitHub↗0
  • deepdoctection/deepdoctectiondeepdoctection avatar

    deepdoctection/deepdoctection

    3,181View on GitHub↗

    A Repo For Document AI

    Python
    View on GitHub↗3,181
See all 16 alternatives to Pdf Text Extraction Benchmark→

Frequently asked questions

What does ckorzen/pdf-text-extraction-benchmark do?

This project is about benchmarking and evaluating existing PDF extraction tools on their semantic abilities to extract the body texts from PDF documents, especially from scientific articles. It provides (1) a benchmark generator, (2) a ready-to-use benchmark and (3) an extensive evaluation, with…

What are the main features of ckorzen/pdf-text-extraction-benchmark?

The main features of ckorzen/pdf-text-extraction-benchmark are: PDF Processing Tools.

What are some open-source alternatives to ckorzen/pdf-text-extraction-benchmark?

Open-source alternatives to ckorzen/pdf-text-extraction-benchmark include: frooodle/stirling-pdf — Stirling-PDF is a web-based PDF management suite used for editing, merging, splitting, and converting PDF documents.… guaguastandup/zotero-pdf2zh — zotero-pdf2zh is a translation system and Zotero plugin designed to convert academic PDF papers into target languages… applicaai/digital-born-pdf-scanner — Many of PDF files that we have downloaded are digital-born, that is contain easily accessible text layer that PDF… deepdoctection/deepdoctection — A Repo For Document AI. apache/pdfbox — Licensed to the Apache Software Foundation (ASF) under one or more contributor license agreements. See the NOTICE file… allenai/pawls — Demo Server | Video Tutorial | Paper.