3 个仓库
Returns precise coordinates for every text line and table cell, preserving document layout for downstream geometric and visual analysis.
Distinct from Spatial Bounding Box Management: Distinct from Spatial Bounding Box Management: focuses on extracting bounding boxes from document layouts, not geospatial clipping or membership tests.
Explore 3 awesome GitHub repositories matching scientific & mathematical computing · Document Layout Bounding Box Extractors. Refine with filters or upvote what's useful.
A fast, helpful, and open-source document parser
Returns precise coordinates for every text line and table cell, preserving document layout for downstream analysis.
Grobid 是一个机器学习系统,旨在将学术和科学 PDF 出版物转换为结构化的 XML。它作为一个 PDF 转 XML 解析器和学术元数据提取器,从研究论文中识别并规范化标题、作者、所属机构和参考文献。 该系统利用深度学习文档分割器将原始 PDF 分割为功能区域,并采用参考文献解析器将引文与外部注册表进行匹配,以进行元数据丰富和 DOI 解析。它支持完整的机器学习模型训练流水线,允许生成标注训练语料库、模型再训练以及导出模型二进制文件。 该项目涵盖了广泛的提取功能,包括文档标题解析、全文正文结构化,以及资助信息和专利引文等领域特定实体的识别。它还提供用于边界框提取和坐标映射的空间分析工具,以将语义标签与原始 PDF 布局同步。 该应用程序可通过容器化镜像部署,并包含用于大型文档集合多线程批处理的命令行工具。
Extracts bounding box coordinates and font styles to improve the accuracy of structural document recognition.
HTML-Renderer is a managed C# library that processes HTML and CSS content to render desktop user interfaces, generate image files, and export PDF documents. Built entirely in managed code without external native dependencies, the library parses markup into a structured document object model, applies cascading style rules, and computes virtual box layouts directly in memory. The rendering pipeline features a direct bitmap rasterisation engine that draws styled layouts straight onto graphics targets and document canvases, along with a dedicated pagination engine that splits continuous layouts a
Computes element dimensions and coordinates entirely in memory to prepare content for drawing surfaces.