awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 repositorios

Awesome GitHub RepositoriesDocument Segmentation

Techniques for decomposing images into hierarchical text structures like lines and characters.

Explore 2 awesome GitHub repositories matching graphics & multimedia · Document Segmentation. Refine with filters or upvote what's useful.

Awesome Document Segmentation GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • tesseract-ocr/tesseractAvatar de tesseract-ocr

    tesseract-ocr/tesseract

    74,751Ver en GitHub↗

    Tesseract is a neural network-based optical character recognition engine designed to convert scanned images and digital documents into machine-readable, searchable text. It functions as both a command-line utility for automating large-scale digitization workflows and a cross-platform library that can be embedded into desktop, mobile, or server-side applications. By utilizing long short-term memory networks, the engine provides robust text extraction across more than one hundred languages and dozens of scripts. The project distinguishes itself through a sophisticated document layout analysis f

    Decomposes visual documents into hierarchical structures, including text blocks, lines, and individual characters.

    C++hacktoberfestlstmmachine-learning
    Ver en GitHub↗74,751
  • datalab-to/suryaAvatar de datalab-to

    datalab-to/surya

    20,889Ver en GitHub↗

    Surya is a document processing platform designed to transform unstructured files into structured, machine-readable data. It provides a comprehensive suite of tools for text recognition, layout analysis, and reading order detection, enabling the conversion of PDFs and images into formats such as JSON, HTML, or markdown. The platform is built to handle complex document workflows, offering capabilities for data extraction, document segmentation, and automated form completion. The platform distinguishes itself through a robust pipeline-based architecture that allows users to chain analysis tasks

    Identifies and isolates distinct sections within documents to improve data extraction accuracy.

    Python
    Ver en GitHub↗20,889
  1. Home
  2. Graphics & Multimedia
  3. Media Processing and Analysis
  4. Media Manipulation
  5. Media Processing Workflows
  6. Computer Vision Pipelines
  7. Document Segmentation