awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

17 repository-uri

Awesome GitHub RepositoriesDocument Ingestion Pipelines

Frameworks for extracting, parsing, and chunking raw documents into vector embeddings for semantic search and retrieval.

Distinguishing note: Focuses on the ETL and vectorization process for unstructured data, distinct from general-purpose database management.

Explore 17 awesome GitHub repositories matching data & databases · Document Ingestion Pipelines. Refine with filters or upvote what's useful.

Awesome Document Ingestion Pipelines GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • quivrhq/quivrAvatar QuivrHQ

    QuivrHQ/quivr

    39,165Vezi pe GitHub↗

    Quivr is a retrieval-augmented generation platform designed to transform raw documents into searchable knowledge bases. It functions as a centralized environment where users can ingest files, index them into vector databases, and interact with language models to receive contextually relevant, data-backed responses. The platform distinguishes itself through an agentic workflow orchestrator that sequences retrieval tasks, tool execution, and model interactions to resolve complex, multi-step queries. This engine is entirely configuration-driven, allowing users to define document ingestion, chunk

    A structured process that handles file parsing, text chunking, and vector embedding management to transform raw documents into searchable knowledge bases.

    Pythonaiapichatbot
    Vezi pe GitHub↗39,165
  • openai/chatgpt-retrieval-pluginAvatar openai

    openai/chatgpt-retrieval-plugin

    21,192Vezi pe GitHub↗

    This project is a retrieval-augmented generation pipeline designed for building custom ChatGPT plugins that allow language models to query private or professional documents. It implements a full retrieval workflow, from processing and indexing document chunks to retrieving relevant context for natural language queries. The system distinguishes itself through a hybrid retrieval approach that combines dense vector embeddings with sparse keyword matching, further refined by a two-stage semantic re-ranking process. It includes specialized data privacy tools for screening personally identifiable i

    Processes JSON document dumps to store content and associated metadata in a vector database.

    Pythonchatgptchatgpt-plugins
    Vezi pe GitHub↗21,192
  • arc53/docsgptAvatar arc53

    arc53/DocsGPT

    17,939Vezi pe GitHub↗

    DocsGPT is a retrieval-augmented generation platform and private knowledge base used to build AI agents that perform grounded search and analysis. It functions as a multi-model AI orchestrator and enterprise agent builder, allowing for the integration of various local and cloud language models to customize reasoning and text generation. The project provides a visual environment for developing automated assistants using conditional logic and third-party API connectivity. It enables the creation of private AI agents capable of performing enterprise search and detailed document analysis using pr

    Implements an asynchronous pipeline for extracting and vectorizing diverse documents into a searchable knowledge base.

    Pythonagent-builderagentsai
    Vezi pe GitHub↗17,939
  • unstructured-io/unstructuredAvatar Unstructured-IO

    Unstructured-IO/unstructured

    14,019Vezi pe GitHub↗

    Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into structured, machine-readable formats. It functions as a comprehensive platform for document ingestion, partitioning, and enrichment, specifically engineered to prepare complex data for retrieval-augmented generation and agentic AI workflows. The platform distinguishes itself through its sophisticated document processing strategies, which combine rule-based extraction with vision-language models to handle diverse file layouts, tables, and images. It provides a modular architecture t

    Connects to content repositories to retrieve documents while capturing associated permission metadata.

    HTMLdata-pipelinesdeep-learningdocument-image-analysis
    Vezi pe GitHub↗14,019
  • spring-projects/spring-aiAvatar spring-projects

    spring-projects/spring-ai

    9,001Vezi pe GitHub↗

    Spring AI is an application framework for Java that provides a portable, fluent API for integrating AI models, tools, and vector stores into applications. It wraps multiple AI providers behind a common interface, allowing developers to switch between chat, embedding, image, and speech models without changing application code. The framework includes a chainable chat client API similar to WebClient or RestClient, supports both synchronous and streaming interactions, and offers structured output conversion that transforms unstructured AI responses into strongly-typed Java objects. The framework

    Extracts, transforms, and loads documents into a vector store for retrieval-augmented generation pipelines.

    Javaartificial-intelligencejavaspring-ai
    Vezi pe GitHub↗9,001
  • runanywhereai/runanywhere-sdksAvatar RunanywhereAI

    RunanywhereAI/runanywhere-sdks

    8,781Vezi pe GitHub↗

    This project is an on-device AI SDK providing a framework for running large language models, vision models, and speech models locally. It serves as an orchestration layer for local LLM execution, ensuring data privacy and offline availability by utilizing hardware acceleration on the device. The SDK is distinguished by its comprehensive voice and multimodal capabilities, including a coordinated voice pipeline for activity detection, speech-to-text, and text-to-speech synthesis. It also provides a dedicated implementation kit for local retrieval-augmented generation and tools for processing co

    Provides a pipeline for chunking, embedding, and indexing raw documents to facilitate local vector search.

    C++androidapple-intelligencecpp
    Vezi pe GitHub↗8,781
  • ymcui/chinese-llama-alpaca-2Avatar ymcui

    ymcui/Chinese-LLaMA-Alpaca-2

    7,136Vezi pe GitHub↗

    This project provides a Chinese large language model based on the LLaMA architecture. It is an instruction-tuned model optimized for natural language processing and multi-turn conversations in Chinese. The system includes a framework for parameter-efficient fine-tuning using low-rank adaptation and quantization to reduce memory requirements. It also implements retrieval augmented generation for local document question answering and supports long-context processing for sequences up to 64K tokens. The project covers a broad set of capabilities including supervised instruction tuning, reinforce

    Processes common file types into a searchable vector store to facilitate retrieval augmented generation.

    Python64kalpacaalpaca-2
    Vezi pe GitHub↗7,136
  • vespa-engine/vespaAvatar vespa-engine

    vespa-engine/vespa

    6,961Vezi pe GitHub↗

    Vespa is a distributed search engine, vector database, and machine learning ranking engine. It serves as an AI search platform designed to handle large-scale document indexing and complex query processing across a cluster of nodes, combining keyword retrieval with high-dimensional embedding storage for semantic similarity search. The platform distinguishes itself by integrating machine learning models directly into the search pipeline to perform real-time inference and ranking. It converts these models into ranking expressions to score and order results based on relevance, while providing a s

    Routes document operations through chainable processors to transform and prepare data before indexing.

    Java
    Vezi pe GitHub↗6,961
  • feast-dev/feastAvatar feast-dev

    feast-dev/feast

    6,727Vezi pe GitHub↗

    Feast is an open-source feature store for machine learning that provides a central platform for defining, storing, and serving features across both training and inference workflows. It operates as a declarative system where feature definitions are written as code in Python files, synchronized to a central registry, and made available for low-latency online retrieval or point-in-time correct historical joins for training datasets. The project abstracts storage behind a pluggable architecture, allowing offline and online backends to be swapped without changing retrieval logic, and coordinates ma

    Chunks, embeds, and writes documents into the feature store in a single configurable pipeline.

    Pythonbig-datadata-engineeringdata-quality
    Vezi pe GitHub↗6,727
  • voltagent/voltagentAvatar VoltAgent

    VoltAgent/voltagent

    6,020Vezi pe GitHub↗

    Ingests documents in multiple formats and automatically chunks, embeds, and indexes them for semantic search.

    TypeScriptagentsaiai-agents
    Vezi pe GitHub↗6,020
  • googlecloudplatform/agent-starter-packAvatar GoogleCloudPlatform

    GoogleCloudPlatform/agent-starter-pack

    5,752Vezi pe GitHub↗

    Ingests documents into search or vector datastores via cloud storage uploads or pipeline runs.

    Pythonagentsgcpgemini
    Vezi pe GitHub↗5,752
  • pyspur-dev/pyspurAvatar PySpur-Dev

    PySpur-Dev/pyspur

    5,677Vezi pe GitHub↗

    Ingests files through a multi-stage pipeline of parsing, chunking, embedding, and upserting into vector indices.

    TypeScriptagentagentsai
    Vezi pe GitHub↗5,677
  • alibaba/zvecAvatar alibaba

    alibaba/zvec

    5,198Vezi pe GitHub↗

    zvec is an embedded vector database engine and indexing library designed for high-dimensional similarity search. It functions as a hybrid search engine and a retrieval-augmented generation knowledge base, allowing for the storage and retrieval of dense and sparse vectors. The system is distinguished by its hybrid retrieval pipeline, which fuses vector similarity, full-text keyword matching, and scalar metadata filtering into single query operations. It supports a plugin-based model integration system for registering custom embedding models and rerankers, as well as language bindings for nativ

    Implements pipelines for storing documents containing both scalar metadata and high-dimensional vector embeddings.

    C++ann-searchembedded-databaserag
    Vezi pe GitHub↗5,198
  • truefoundry/cognitaAvatar truefoundry

    truefoundry/cognita

    4,317Vezi pe GitHub↗

    Cognita is a retrieval augmented generation orchestration framework used to build pipelines that connect document stores and language models to provide grounded answers. It functions as a document ingestion pipeline and a vector database integrator, managing the process of loading, parsing, and indexing files into a searchable knowledge base. The system includes a language model gateway proxy that provides a unified API to interact with multiple different model providers. This routing layer decouples the application from specific vendors, allowing requests to be proxied through a provider-agn

    Extracts, parses, and chunks raw documents into vector embeddings for semantic search and retrieval.

    Pythonagentaiapplication
    Vezi pe GitHub↗4,317
  • run-llama/llama_cloud_servicesAvatar run-llama

    run-llama/llama_cloud_services

    4,251Vezi pe GitHub↗

    Llama Cloud Services este o platformă de gestionare a cunoștințelor și un serviciu găzduit conceput pentru a parsa, ingera și indexa documente complexe. Funcționează ca o bază de cunoștințe în cloud și un pipeline automat de ingestie care convertește documentele nestructurate în indici căutabili pentru generarea augmentată prin recuperare (RAG). Sistemul folosește agenți autonomi pentru a efectua extracția agentică a datelor, transformând informațiile nestructurate în formate de date structurate. Oferă instrumente pentru administrarea bazei de cunoștințe în cloud, permițând gestionarea repository-urilor găzduite care alimentează agenți specializați de tip large language model. Platforma acoperă o gamă largă de capabilități, inclusiv parsarea complexă a documentelor, ingestia de documente la nivel enterprise și organizarea magazinelor de date bazate pe cloud pentru a oferi context pentru promptarea modelelor.

    Provides a sequential pipeline for parsing and indexing large volumes of documents for searchable retrieval.

    TypeScriptdocumentdocument-parserdocument-parsing
    Vezi pe GitHub↗4,251
  • langflow-ai/openragAvatar langflow-ai

    langflow-ai/openrag

    4,255Vezi pe GitHub↗

    OpenRAG este un framework agentic de generare augmentată prin recuperare (RAG) și un stack containerizat. Acesta oferă un motor de căutare vectorială pentru indexarea documentelor nestructurate și un server Model Context Protocol care expune instrumente de ingestie și căutare semantică către asistenții AI externi. Sistemul se distinge printr-o interfață vizuală de orchestrare AI, permițând utilizatorilor să construiască conducte de recuperare printr-un designer drag-and-drop în loc de cod manual. Utilizează fluxuri de lucru agentice care coordonează mai mulți agenți și pași de re-ranking pentru a îmbunătăți acuratețea răspunsurilor și permite definirea abilităților agenților printr-un format markdown standardizat. Platforma acoperă conducte cuprinzătoare de ingestie a documentelor pentru a parsa date nestructurate, capabilități de căutare semantică enterprise și opțiuni de implementare containerizată cu suport pentru accelerare GPU. De asemenea, include autentificarea interfeței serverului și sincronizarea rolurilor utilizatorilor pentru controlul accesului.

    Parses unstructured real-world data into a searchable format for use in retrieval-augmented generation pipelines.

    Python
    Vezi pe GitHub↗4,255
  • i-am-bee/beeai-frameworkAvatar i-am-bee

    i-am-bee/beeai-framework

    3,304Vezi pe GitHub↗

    The BeeAI Framework is an LLM agent framework and multi-agent orchestration engine used to build autonomous agents that coordinate reasoning, tool execution, and complex workflows. It functions as a structured AI output controller and RAG integration library, providing a unified interface to manage multiple language model providers. The framework is distinguished by its implementation of the Model Context Protocol, allowing agents, tools, and models to be shared between different AI platforms and hosted as agentic tooling servers. It enables the design of collaborative agent teams through dec

    Ships pipelines for extracting, parsing, and chunking raw documents into vector embeddings for semantic search.

    Pythonagentsaiai-agent
    Vezi pe GitHub↗3,304
  1. Home
  2. Data & Databases
  3. Document Ingestion Pipelines

Explorează sub-etichetele

  • Feature Store Document IngestionChunks, embeds, and writes documents to the online feature store in a single step using a configurable pipeline. **Distinct from Document Ingestion Pipelines:** Distinct from Document Ingestion Pipelines: focuses on ingesting documents specifically into a feature store for ML serving, not general document processing.
  • FileNet Connectors1 sub-tagIntegrations for retrieving documents and associated access control metadata from content repositories. **Distinct from Document Ingestion Pipelines:** Focuses on FileNet-specific repository ingestion, distinct from general document ingestion pipelines.