awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to ibm-aur-nlp/pubtabnet

Projects sharing features with PubTabNet

10 open-source projects similar to ibm-aur-nlp/pubtabnet, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • atlanhq/camelotatlanhq avatar

    atlanhq/camelot

    3,717View on GitHub↗

    Camelot is a Python-based library designed to parse, extract, and clean tabular data from PDF files. It converts table elements from text-based PDF documents into programmable data structures and dataframes. The tool identifies tabular regions using coordinate-based grouping, lattice-based line detection, and stream-based text extraction. It can also rasterize PDF pages into images to utilize computer vision for detecting structural lines and boundaries. Extracted data is validated through accuracy and whitespace metrics to filter out low-quality extractions. The processed information can be

    Python
    View on GitHub↗3,717
  • carefree0910/carefree-learncarefree0910 avatar

    carefree0910/carefree-learn

    410View on GitHub↗

    Deep Learning ❤️ PyTorch

    Python
    View on GitHub↗410
  • chezou/tabula-pychezou avatar

    chezou/tabula-py

    2,315View on GitHub↗

    Simple wrapper of tabula-java: extract table from PDF into pandas DataFrame

    Pythonpandaspdfpython
    View on GitHub↗2,315
  • chineseocr/table-ocrchineseocr avatar

    chineseocr/table-ocr

    605View on GitHub↗

    x 支持GPU,CPU(opencv dnn加速); - 整合darknet-ocr完成对表格的重建,输出json\excel

    Python
    View on GitHub↗605
  • diyago/gan-for-tabular-dataDiyago avatar

    Diyago/GAN-for-tabular-data

    0View on GitHub↗

    Generative Networks are well-known for their success in realistic image generation. However, they can also be applied to generate tabular data. We introduce major improvements for generating high-fidelity tabular data giving oppotunity to try GANS, TimeGANs, Diffusions and LLM for tabular data…

    View on GitHub↗0

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • google-research/tapasgoogle-research avatar

    google-research/tapas

    1,203View on GitHub↗

    End-to-end neural table-text understanding models.

    Python
    View on GitHub↗1,203
  • holms-ur/fine-tuningholms-ur avatar

    holms-ur/fine-tuning

    72View on GitHub↗

    Close-Domain fine-tuning for table detection

    Python
    View on GitHub↗72
  • jsvine/pdfplumberjsvine avatar

    jsvine/pdfplumber

    9,732View on GitHub↗

    pdfplumber is a PDF data extraction library and layout analysis tool used to retrieve text, tables, and geometric objects from PDF files using precise coordinate-based analysis. It functions as a layout analyzer and table parser that identifies the bounding boxes and visual coordinates for every character and image on a page. The library distinguishes itself through visual debugging capabilities, allowing users to render PDF pages as images and draw annotations to verify the position of extracted data. It employs line and intersection analysis to identify cell structures and convert unstructu

    Pythonpdfpdf-parsingtable-extraction
    View on GitHub↗9,732
  • paperswithcode/axcellpaperswithcode avatar

    paperswithcode/axcell

    440View on GitHub↗

    Tools for extracting tables and results from Machine Learning papers

    Python
    View on GitHub↗440
  • wzbsocialsciencecenter/pdftabextractWZBSocialScienceCenter avatar

    WZBSocialScienceCenter/pdftabextract

    2,255View on GitHub↗

    A set of tools for extracting tables from PDF files helping to do data mining on (OCR-processed) scanned documents.

    Python
    View on GitHub↗2,255