How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
PubTabNet is a large dataset for image-based table recognition, containing 568k+ images of tabular data annotated with the corresponding HTML representation of the tables. The table images are extracted from the scientific publications included in the PubMed Central Open Access Subset…
The main features of ibm-aur-nlp/pubtabnet are: Table Processing.
Projects with overlapping indexed features include: atlanhq/camelot — Camelot is a Python-based library designed to parse, extract, and clean tabular data from PDF files. It converts table… carefree0910/carefree-learn — Deep Learning ❤️ PyTorch. chezou/tabula-py — Simple wrapper of tabula-java: extract table from PDF into pandas DataFrame. chineseocr/table-ocr — [x] 支持GPU,CPU(opencv dnn加速); - [ ] 整合darknet-ocr完成对表格的重建,输出json\excel. diyago/gan-for-tabular-data — Generative Networks are well-known for their success in realistic image generation. However, they can also be applied… google-research/tapas — End-to-end neural table-text understanding models.
Simple wrapper of tabula-java: extract table from PDF into pandas DataFrame
x 支持GPU,CPU(opencv dnn加速); - 整合darknet-ocr完成对表格的重建,输出json\excel
Camelot is a Python-based library designed to parse, extract, and clean tabular data from PDF files. It converts table elements from text-based PDF documents into programmable data structures and dataframes. The tool identifies tabular regions using coordinate-based grouping, lattice-based line detection, and stream-based text extraction. It can also rasterize PDF pages into images to utilize computer vision for detecting structural lines and boundaries. Extracted data is validated through accuracy and whitespace metrics to filter out low-quality extractions. The processed information can be