8 repository-uri
Workflows that parse raw files into structured text chunks and metadata to facilitate semantic search and data retrieval.
Explore 8 awesome GitHub repositories matching data & databases · Document Ingestion Pipelines. Refine with filters or upvote what's useful.
This platform serves as a comprehensive environment for managing private language models, document knowledge bases, and automated agent workflows within secure local infrastructure. It functions as a document-aware workspace that enables users to ingest diverse file formats into searchable repositories, ensuring that all data processing and model inference remain within private, local environments to maintain data sovereignty. The system distinguishes itself through a modular agentic engine that allows for the definition of custom skills and external tool execution. By utilizing a multi-model
Automates the ingestion of raw documents into structured text and metadata to enable efficient semantic search.
This project is a privacy-first backend service designed to facilitate retrieval-augmented generation by processing local documents into searchable vector representations. It provides a modular architecture that allows users to ingest diverse file formats, manage document metadata, and perform semantic searches to provide context-aware responses for chat and completion requests. The system distinguishes itself through a database-agnostic abstraction layer that supports various storage backends, ranging from local disk storage to enterprise-grade vector databases. It offers flexible deployment
Parses raw files into structured text chunks and metadata to facilitate semantic search and data retrieval.
DeepCode is an agentic development framework designed to orchestrate autonomous AI agents for software engineering tasks. It functions as a multi-agent workflow orchestrator that translates natural language requirements into functional codebases by coordinating specialized agents for architectural planning, intent analysis, and implementation. The platform integrates multiple language models to power these automated routines, providing a unified environment for complex development projects. The system distinguishes itself through its ability to transform academic research papers into executab
Parses large technical research papers into structured text chunks to maintain semantic integrity for language model processing.
Unstructured is an enterprise-grade data orchestration engine designed to transform raw, unstructured files into structured, machine-readable formats. It functions as a comprehensive platform for document ingestion, partitioning, and enrichment, specifically engineered to prepare complex data for retrieval-augmented generation and agentic AI workflows. The platform distinguishes itself through its sophisticated document processing strategies, which combine rule-based extraction with vision-language models to handle diverse file layouts, tables, and images. It provides a modular architecture t
Connects to cloud storage to retrieve documents and process them into structured formats for downstream AI applications.
Paperless is a self-hosted document management system designed to digitize, index, and archive paper documents. It functions as an optical character recognition system that converts scanned images and PDFs into a searchable digital library, providing a web-based interface for querying and retrieving documents from a database. The system features an automated file ingestion pipeline that monitors specific directories and email inboxes to process and import documents without manual uploading. To maintain a private archive, it includes on-disk encryption for sensitive files and the ability to or
Ships a workflow that monitors directories and email to automatically import and process scanned documents.
Acest proiect este un SDK de dezvoltare software și un instrument de gestionare a clusterelor pentru PHP. Servește ca un SDK de căutare full-text și interfață de căutare vectorială, permițând aplicațiilor să efectueze căutări lexicale, fuzzy și semantice asupra datelor indexate. Biblioteca implementează un client HTTP PSR 7 pentru a asigura compatibilitatea cross-environment prin interfețe de mesagerie standardizate. Oferă o interfață specializată pentru recuperarea embedding-urilor și executarea fluxurilor de lucru de recuperare semantică folosind date vectoriale. Suprafața sa de capabilități acoperă o gamă largă de sarcini administrative și operaționale, inclusiv gestionarea indicilor de căutare, monitorizarea stării clusterului și operațiuni privind ciclul de viață al documentelor. Suportă metode diverse de interogare precum SQL, EQL și ES|QL, alături de agregarea datelor și analiza geospațială. În plus, oferă instrumente pentru orchestrarea machine learning-ului, detectarea anomaliilor și gestionarea identității și a accesului.
Enables the creation and simulation of pipelines to process and transform documents before indexing.
Cognita is a retrieval augmented generation orchestration framework used to build pipelines that connect document stores and language models to provide grounded answers. It functions as a document ingestion pipeline and a vector database integrator, managing the process of loading, parsing, and indexing files into a searchable knowledge base. The system includes a language model gateway proxy that provides a unified API to interact with multiple different model providers. This routing layer decouples the application from specific vendors, allowing requests to be proxied through a provider-agn
Ships workflows that parse raw files into structured text chunks and metadata for semantic search.
NeMo-Retriever is a framework designed for building end-to-end document ingestion and retrieval-augmented generation pipelines. It provides a suite of tools for processing, classifying, and structuring diverse file formats, transforming raw enterprise data into searchable information assets for generative artificial intelligence applications. The system distinguishes itself through its specialized capabilities for parsing complex document layouts, including tables, charts, and infographics, using integrated optical character recognition and multi-modal extraction. It utilizes a microservice-b
Processes directories of files through configurable pipelines that split, chunk, and enrich metadata for retrieval systems.