2 个仓库
Freely available document parsers that extract text, tables, and layout from PDFs and office files into structured formats.
Distinct from Documentation Parsers: Distinct from Documentation Parsers: focuses on parsing document files (PDFs, office docs) for content extraction, not source code documentation extraction.
Explore 2 awesome GitHub repositories matching programming languages & runtimes · Open-Source Document Parsers. Refine with filters or upvote what's useful.
A fast, helpful, and open-source document parser
An open-source document parser that extracts text, tables, and layout from PDFs and office files into Markdown or JSON.
Textract 是一个多格式文本提取工具和解析器。它提供了一个统一的接口,用于从各种来源(包括文档、图像和音频文件)中提取纯文本。 该系统作为 PDF 和电子表格的文档内容解析器、使用光学字符识别 (OCR) 的图像文本提取器,以及音频录音的语音转文本转录器。
Maps PDFs and spreadsheets to a common internal representation for consistent text retrieval.