2 रिपॉजिटरी
Freely available document parsers that extract text, tables, and layout from PDFs and office files into structured formats.
Distinct from Documentation Parsers: Distinct from Documentation Parsers: focuses on parsing document files (PDFs, office docs) for content extraction, not source code documentation extraction.
Explore 2 awesome GitHub repositories matching programming languages & runtimes · Open-Source Document Parsers. Refine with filters or upvote what's useful.
A fast, helpful, and open-source document parser
An open-source document parser that extracts text, tables, and layout from PDFs and office files into Markdown or JSON.
Textract एक मल्टी-फॉर्मेट टेक्स्ट एक्सट्रैक्शन टूल और पार्सर है। यह दस्तावेज़ों, इमेजेस और ऑडियो फ़ाइलों सहित विभिन्न स्रोतों से प्लेन टेक्स्ट निकालने के लिए एक एकीकृत इंटरफेस प्रदान करता है। यह सिस्टम PDF और स्प्रेडशीट्स के लिए एक दस्तावेज़ कंटेंट पार्सर, ऑप्टिकल कैरेक्टर रिकग्निशन का उपयोग करने वाला एक इमेज टेक्स्ट एक्सट्रैक्टर, और ऑडियो रिकॉर्डिंग के लिए स्पीच-टू-टेक्स्ट ट्रांसक्राइबर के रूप में कार्य करता है।
Maps PDFs and spreadsheets to a common internal representation for consistent text retrieval.