awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 रिपॉजिटरी

Awesome GitHub RepositoriesVision-Language Model Backends

Uses vision-language models as an OCR backend and for extracting structured JSON from documents using a schema.

Distinct from Structured Document Extraction: Distinct from Structured Document Extraction: focuses on using VLM backends for extraction, not general layout-to-text conversion.

Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Vision-Language Model Backends. Refine with filters or upvote what's useful.

Awesome Vision-Language Model Backends GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • kreuzberg-dev/kreuzbergkreuzberg-dev का अवतार

    kreuzberg-dev/kreuzberg

    8,527GitHub पर देखें↗

    Kreuzberg is a document extraction engine that converts PDFs, Office files, images, and over 90 other formats into clean, structured text and metadata. It is built around a compiled Rust core that can be used as a native library, a command-line tool, a REST API server, or a WebAssembly module for browser-based processing. The system is designed to run entirely on self-hosted infrastructure, with no data leaving the user's environment. What distinguishes Kreuzberg is its breadth of integration surfaces and its pipeline architecture. It exposes extraction capabilities through native bindings fo

    Leverages vision language models as an OCR backend and extracts structured JSON from documents using a schema, supporting 146 LLM providers.

    Rustdocument-intelligenceelixirffi
    GitHub पर देखें↗8,527
  • katanaml/sparrowkatanaml का अवतार

    katanaml/sparrow

    5,162GitHub पर देखें↗

    Sparrow एक LLM डॉक्यूमेंट एक्सट्रैक्शन प्लेटफॉर्म और विज़न-आधारित इन्फरेंस इंजन है जिसे छवियों और PDFs को वैलिडेटेड स्ट्रक्चर्ड डेटा में बदलने के लिए डिज़ाइन किया गया है। यह एक एजेंटिक वर्कफ़्लो ऑर्केस्ट्रेटर के रूप में कार्य करता है जो वर्गीकरण, निष्कर्षण और सत्यापन कार्यों को मल्टी-स्टेप पाइपलाइनों में जोड़ता है। यह सिस्टम एक बैकएंड-अज्ञेयवादी इन्फरेंस लेयर के माध्यम से खुद को अलग करता है जो लोकल GPUs, Apple Silicon और क्लाउड प्रदाताओं के बीच मॉडल का प्रबंधन करता है। यह एक्सट्रैक्ट किए गए टेक्स्ट को सटीक बाउंडिंग बॉक्स निर्देशांकों पर मैप करने के लिए कोऑर्डिनेट-आधारित विजुअल ग्राउंडिंग का उपयोग करता है और ध्यान केंद्रित करने और डेटा फॉर्मेट को सामान्य करने के लिए हिंट-आधारित मॉडल स्टीयरिंग का उपयोग करता है। यह प्लेटफॉर्म डॉक्यूमेंट इंटेलिजेंस वर्कफ़्लो को कवर करता है, जिसमें संरचनात्मक अखंडता बनाए रखने के लिए विशेष इमेज-आधारित टेबल प्रोसेसिंग और एक्सट्रैक्ट किए गए फील्ड्स की शुद्धता को सत्यापित करने के लिए स्कीमा-आधारित सत्यापन शामिल है। यह API प्रदर्शन, उपयोग एनालिटिक्स और सिस्टम हेल्थ की निगरानी के लिए एक डॉक्यूमेंट एनालिसिस डैशबोर्ड भी प्रदान करता है। आर्किटेक्चर में इंडेक्सिंग और ऑर्केस्ट्रेशन के लिए उपयोग की जाने वाली थर्ड-पार्टी लाइब्रेरीज़ को एकीकृत करने के लिए एक प्लगइन-आधारित एक्सटेंशन सिस्टम शामिल है।

    Uses vision-capable language models to parse document layouts and convert visual content into structured data.

    Pythonagentic-aicomputer-visiondocumentai
    GitHub पर देखें↗5,162
  1. Home
  2. Artificial Intelligence & ML
  3. Natural Language Processing
  4. Structured Document Extraction
  5. Vision-Language Model Backends