2 Repos
Freely available document parsers that extract text, tables, and layout from PDFs and office files into structured formats.
Distinct from Documentation Parsers: Distinct from Documentation Parsers: focuses on parsing document files (PDFs, office docs) for content extraction, not source code documentation extraction.
Explore 2 awesome GitHub repositories matching programming languages & runtimes · Open-Source Document Parsers. Refine with filters or upvote what's useful.
A fast, helpful, and open-source document parser
An open-source document parser that extracts text, tables, and layout from PDFs and office files into Markdown or JSON.
Textract ist ein Tool und Parser zur Textextraktion aus verschiedenen Formaten. Es bietet eine einheitliche Schnittstelle, um Klartext aus einer Vielzahl von Quellen zu extrahieren, einschließlich Dokumenten, Bildern und Audiodateien. Das System fungiert als Dokumenten-Content-Parser für PDFs und Tabellenkalkulationen, als Bildtextextraktor mittels optischer Zeichenerkennung (OCR) und als Speech-to-Text-Transkribierer für Audioaufnahmen.
Maps PDFs and spreadsheets to a common internal representation for consistent text retrieval.