awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
A

applicaai/digital-born-pdf-scanner

0
View on GitHub↗
0 stars·0 forks·12 views

Digital Born Pdf Scanner

Many of PDF files that we have downloaded are digital-born, that is contain easily accessible text layer that PDF viewers use to display text. Some are definitely scanned documents, that do not have any text layer at all, some are searchable OCR-processed scans that contain a lot of hidden text.

Features

  • PDF Processing Tools - Utility to verify if a PDF is born-digital.

Star history

Star history chart for applicaai/digital-born-pdf-scannerStar history chart for applicaai/digital-born-pdf-scanner

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Digital Born Pdf Scanner

These projects share indexed features with Digital Born Pdf Scanner. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • frooodle/stirling-pdfFrooodle avatar

    Frooodle/Stirling-PDF

    81,168View on GitHub↗

    Stirling-PDF is a web-based PDF management suite used for editing, merging, splitting, and converting PDF documents. It functions as a self-hosted document manager, providing a centralized interface for users to manipulate files on a private server. The system features a workflow automation engine that allows for the creation of processing pipelines to handle large volumes of documents without writing custom code. It also includes an optical character recognition tool to convert scanned PDFs into searchable and editable text. Access is managed through single sign-on integration and OIDC comp

    Java
    View on GitHub↗81,168
  • guaguastandup/zotero-pdf2zhguaguastandup avatar

    guaguastandup/zotero-pdf2zh

    3,008View on GitHub↗

    zotero-pdf2zh is a translation system and Zotero plugin designed to convert academic PDF papers into target languages using large language model APIs and machine translation services. It functions as an LLM document translator that integrates directly into Zotero libraries to process research documents. The tool distinguishes itself by generating bilingual PDFs with parallel layouts for side-by-side comparison of original and translated text. It also includes a mobile layout optimizer that crops double-column PDFs and stacks them vertically to improve readability on narrow screens. The syste

    Python
    View on GitHub↗3,008
  • ckorzen/pdf-text-extraction-benchmarkC

    ckorzen/pdf-text-extraction-benchmark

    0View on GitHub↗

    This project is about benchmarking and evaluating existing PDF extraction tools on their semantic abilities to extract the body texts from PDF documents, especially from scientific articles. It provides (1) a benchmark generator, (2) a ready-to-use benchmark and (3) an extensive evaluation, with…

    View on GitHub↗0
  • deepdoctection/deepdoctectiondeepdoctection avatar

    deepdoctection/deepdoctection

    3,181View on GitHub↗

    A Repo For Document AI

    Python
    View on GitHub↗3,181
Compare all 16 related projects→

Frequently asked questions

What does applicaai/digital-born-pdf-scanner do?

Many of PDF files that we have downloaded are digital-born, that is contain easily accessible text layer that PDF viewers use to display text. Some are definitely scanned documents, that do not have any text layer at all, some are searchable OCR-processed scans that contain a lot of hidden text.

What are the main features of applicaai/digital-born-pdf-scanner?

The main features of applicaai/digital-born-pdf-scanner are: PDF Processing Tools.

Which projects share features with applicaai/digital-born-pdf-scanner?

Projects with overlapping indexed features include: frooodle/stirling-pdf — Stirling-PDF is a web-based PDF management suite used for editing, merging, splitting, and converting PDF documents.… guaguastandup/zotero-pdf2zh — zotero-pdf2zh is a translation system and Zotero plugin designed to convert academic PDF papers into target languages… ckorzen/pdf-text-extraction-benchmark — This project is about benchmarking and evaluating existing PDF extraction tools on their semantic abilities to extract… deepdoctection/deepdoctection — A Repo For Document AI. apache/pdfbox — Licensed to the Apache Software Foundation (ASF) under one or more contributor license agreements. See the NOTICE file… allenai/pawls — Demo Server | Video Tutorial | Paper.