awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
ibm-aur-nlp avatar

ibm-aur-nlp/PubTabNet

0
View on GitHub↗
483 stars·86 forks·Jupyter Notebook·14 views

PubTabNet

PubTabNet is a large dataset for image-based table recognition, containing 568k+ images of tabular data annotated with the corresponding HTML representation of the tables. The table images are extracted from the scientific publications included in the PubMed Central Open Access Subset…

Features

  • Table Processing - Model for segmenting and identifying tables in documents.

Star history

Star history chart for ibm-aur-nlp/pubtabnetStar history chart for ibm-aur-nlp/pubtabnet

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does ibm-aur-nlp/pubtabnet do?

PubTabNet is a large dataset for image-based table recognition, containing 568k+ images of tabular data annotated with the corresponding HTML representation of the tables. The table images are extracted from the scientific publications included in the PubMed Central Open Access Subset…

What are the main features of ibm-aur-nlp/pubtabnet?

The main features of ibm-aur-nlp/pubtabnet are: Table Processing.

Which projects share features with ibm-aur-nlp/pubtabnet?

Projects with overlapping indexed features include: atlanhq/camelot — Camelot is a Python-based library designed to parse, extract, and clean tabular data from PDF files. It converts table… carefree0910/carefree-learn — Deep Learning ❤️ PyTorch. chezou/tabula-py — Simple wrapper of tabula-java: extract table from PDF into pandas DataFrame. chineseocr/table-ocr — [x] 支持GPU,CPU(opencv dnn加速); - [ ] 整合darknet-ocr完成对表格的重建,输出json\excel. diyago/gan-for-tabular-data — Generative Networks are well-known for their success in realistic image generation. However, they can also be applied… google-research/tapas — End-to-end neural table-text understanding models.

Projects sharing features with PubTabNet

These projects share indexed features with PubTabNet. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • carefree0910/carefree-learncarefree0910 avatar

    carefree0910/carefree-learn

    410View on GitHub↗

    Deep Learning ❤️ PyTorch

    Python
    View on GitHub↗410
  • chezou/tabula-pychezou avatar

    chezou/tabula-py

    2,315View on GitHub↗

    Simple wrapper of tabula-java: extract table from PDF into pandas DataFrame

    Pythonpandaspdfpython
    View on GitHub↗2,315
  • chineseocr/table-ocrchineseocr avatar

    chineseocr/table-ocr

    605View on GitHub↗

    x 支持GPU,CPU(opencv dnn加速); - 整合darknet-ocr完成对表格的重建,输出json\excel

    Python
    View on GitHub↗605
  • atlanhq/camelotatlanhq avatar

    atlanhq/camelot

    3,717View on GitHub↗

    Camelot is a Python-based library designed to parse, extract, and clean tabular data from PDF files. It converts table elements from text-based PDF documents into programmable data structures and dataframes. The tool identifies tabular regions using coordinate-based grouping, lattice-based line detection, and stream-based text extraction. It can also rasterize PDF pages into images to utilize computer vision for detecting structural lines and boundaries. Extracted data is validated through accuracy and whitespace metrics to filter out low-quality extractions. The processed information can be

    Python
    View on GitHub↗3,717
Compare all 10 related projects→