17 repository-uri
Frameworks for extracting, parsing, and chunking raw documents into vector embeddings for semantic search and retrieval.
Distinguishing note: Focuses on the ETL and vectorization process for unstructured data, distinct from general-purpose database management.
Explore 17 awesome GitHub repositories matching data & databases · Document Ingestion Pipelines. Refine with filters or upvote what's useful.
Quivr is a retrieval-augmented generation platform designed to transform raw documents into searchable knowledge bases. It functions as a centralized environment where users can ingest files, index them into vector databases, and interact with language models to receive contextually relevant, data-backed responses. The platform distinguishes itself through an agentic workflow orchestrator that sequences retrieval tasks, tool execution, and model interactions to resolve complex, multi-step queries. This engine is entirely configuration-driven, allowing users to define document ingestion, chunk
A structured process that handles file parsing, text chunking, and vector embedding management to transform raw documents into searchable knowledge bases.
This project is a retrieval-augmented generation pipeline designed for building custom ChatGPT plugins that allow language models to query private or professional documents. It implements a full retrieval workflow, from processing and indexing document chunks to retrieving relevant context for natural language queries. The system distinguishes itself through a hybrid retrieval approach that combines dense vector embeddings with sparse keyword matching, further refined by a two-stage semantic re-ranking process. It includes specialized data privacy tools for screening personally identifiable i
Processes JSON document dumps to store content and associated metadata in a vector database.
DocsGPT is a retrieval-augmented generation platform and private knowledge base used to build AI agents that perform grounded search and analysis. It functions as a multi-model AI orchestrator and enterprise agent builder, allowing for the integration of various local and cloud language models to customize reasoning and text generation. The project provides a visual environment for developing automated assistants using conditional logic and third-party API connectivity. It enables the creation of private AI agents capable of performing enterprise search and detailed document analysis using pr
Implements an asynchronous pipeline for extracting and vectorizing diverse documents into a searchable knowledge base.
Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into structured, machine-readable formats. It functions as a comprehensive platform for document ingestion, partitioning, and enrichment, specifically engineered to prepare complex data for retrieval-augmented generation and agentic AI workflows. The platform distinguishes itself through its sophisticated document processing strategies, which combine rule-based extraction with vision-language models to handle diverse file layouts, tables, and images. It provides a modular architecture t
Connects to content repositories to retrieve documents while capturing associated permission metadata.
Spring AI is an application framework for Java that provides a portable, fluent API for integrating AI models, tools, and vector stores into applications. It wraps multiple AI providers behind a common interface, allowing developers to switch between chat, embedding, image, and speech models without changing application code. The framework includes a chainable chat client API similar to WebClient or RestClient, supports both synchronous and streaming interactions, and offers structured output conversion that transforms unstructured AI responses into strongly-typed Java objects. The framework
Extracts, transforms, and loads documents into a vector store for retrieval-augmented generation pipelines.
This project is an on-device AI SDK providing a framework for running large language models, vision models, and speech models locally. It serves as an orchestration layer for local LLM execution, ensuring data privacy and offline availability by utilizing hardware acceleration on the device. The SDK is distinguished by its comprehensive voice and multimodal capabilities, including a coordinated voice pipeline for activity detection, speech-to-text, and text-to-speech synthesis. It also provides a dedicated implementation kit for local retrieval-augmented generation and tools for processing co
Provides a pipeline for chunking, embedding, and indexing raw documents to facilitate local vector search.
This project provides a Chinese large language model based on the LLaMA architecture. It is an instruction-tuned model optimized for natural language processing and multi-turn conversations in Chinese. The system includes a framework for parameter-efficient fine-tuning using low-rank adaptation and quantization to reduce memory requirements. It also implements retrieval augmented generation for local document question answering and supports long-context processing for sequences up to 64K tokens. The project covers a broad set of capabilities including supervised instruction tuning, reinforce
Processes common file types into a searchable vector store to facilitate retrieval augmented generation.
Vespa is a distributed search engine, vector database, and machine learning ranking engine. It serves as an AI search platform designed to handle large-scale document indexing and complex query processing across a cluster of nodes, combining keyword retrieval with high-dimensional embedding storage for semantic similarity search. The platform distinguishes itself by integrating machine learning models directly into the search pipeline to perform real-time inference and ranking. It converts these models into ranking expressions to score and order results based on relevance, while providing a s
Routes document operations through chainable processors to transform and prepare data before indexing.
Feast is an open-source feature store for machine learning that provides a central platform for defining, storing, and serving features across both training and inference workflows. It operates as a declarative system where feature definitions are written as code in Python files, synchronized to a central registry, and made available for low-latency online retrieval or point-in-time correct historical joins for training datasets. The project abstracts storage behind a pluggable architecture, allowing offline and online backends to be swapped without changing retrieval logic, and coordinates ma
Chunks, embeds, and writes documents into the feature store in a single configurable pipeline.
Ingests documents in multiple formats and automatically chunks, embeds, and indexes them for semantic search.
Ingests documents into search or vector datastores via cloud storage uploads or pipeline runs.
Ingests files through a multi-stage pipeline of parsing, chunking, embedding, and upserting into vector indices.
zvec is an embedded vector database engine and indexing library designed for high-dimensional similarity search. It functions as a hybrid search engine and a retrieval-augmented generation knowledge base, allowing for the storage and retrieval of dense and sparse vectors. The system is distinguished by its hybrid retrieval pipeline, which fuses vector similarity, full-text keyword matching, and scalar metadata filtering into single query operations. It supports a plugin-based model integration system for registering custom embedding models and rerankers, as well as language bindings for nativ
Implements pipelines for storing documents containing both scalar metadata and high-dimensional vector embeddings.
Cognita is a retrieval augmented generation orchestration framework used to build pipelines that connect document stores and language models to provide grounded answers. It functions as a document ingestion pipeline and a vector database integrator, managing the process of loading, parsing, and indexing files into a searchable knowledge base. The system includes a language model gateway proxy that provides a unified API to interact with multiple different model providers. This routing layer decouples the application from specific vendors, allowing requests to be proxied through a provider-agn
Extracts, parses, and chunks raw documents into vector embeddings for semantic search and retrieval.
Llama Cloud Services este o platformă de gestionare a cunoștințelor și un serviciu găzduit conceput pentru a parsa, ingera și indexa documente complexe. Funcționează ca o bază de cunoștințe în cloud și un pipeline automat de ingestie care convertește documentele nestructurate în indici căutabili pentru generarea augmentată prin recuperare (RAG). Sistemul folosește agenți autonomi pentru a efectua extracția agentică a datelor, transformând informațiile nestructurate în formate de date structurate. Oferă instrumente pentru administrarea bazei de cunoștințe în cloud, permițând gestionarea repository-urilor găzduite care alimentează agenți specializați de tip large language model. Platforma acoperă o gamă largă de capabilități, inclusiv parsarea complexă a documentelor, ingestia de documente la nivel enterprise și organizarea magazinelor de date bazate pe cloud pentru a oferi context pentru promptarea modelelor.
Provides a sequential pipeline for parsing and indexing large volumes of documents for searchable retrieval.
OpenRAG este un framework agentic de generare augmentată prin recuperare (RAG) și un stack containerizat. Acesta oferă un motor de căutare vectorială pentru indexarea documentelor nestructurate și un server Model Context Protocol care expune instrumente de ingestie și căutare semantică către asistenții AI externi. Sistemul se distinge printr-o interfață vizuală de orchestrare AI, permițând utilizatorilor să construiască conducte de recuperare printr-un designer drag-and-drop în loc de cod manual. Utilizează fluxuri de lucru agentice care coordonează mai mulți agenți și pași de re-ranking pentru a îmbunătăți acuratețea răspunsurilor și permite definirea abilităților agenților printr-un format markdown standardizat. Platforma acoperă conducte cuprinzătoare de ingestie a documentelor pentru a parsa date nestructurate, capabilități de căutare semantică enterprise și opțiuni de implementare containerizată cu suport pentru accelerare GPU. De asemenea, include autentificarea interfeței serverului și sincronizarea rolurilor utilizatorilor pentru controlul accesului.
Parses unstructured real-world data into a searchable format for use in retrieval-augmented generation pipelines.
The BeeAI Framework is an LLM agent framework and multi-agent orchestration engine used to build autonomous agents that coordinate reasoning, tool execution, and complex workflows. It functions as a structured AI output controller and RAG integration library, providing a unified interface to manage multiple language model providers. The framework is distinguished by its implementation of the Model Context Protocol, allowing agents, tools, and models to be shared between different AI platforms and hosted as agentic tooling servers. It enables the design of collaborative agent teams through dec
Ships pipelines for extracting, parsing, and chunking raw documents into vector embeddings for semantic search.