awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
cocoindex-io avatar

cocoindex-io/cocoindex

0
View on GitHub↗
cocoindex.io↗

Cocoindex

Cocoindex is an incremental data processing engine that builds and maintains live indexes for AI agents, with a core focus on codebase indexing and knowledge graph extraction. The engine uses a function-graph execution model where user-defined Python functions are composed into a directed acyclic graph, and it processes data incrementally so only changed source records or code paths are re-computed, avoiding full recomputation at any scale. It supports automatic schema inference from transformation pipeline type annotations and provides full data lineage tracing, tagging every output record with its source items and transformation version.

The project distinguishes itself through declarative target-state reconciliation, where users describe the desired end state of a data store in Python and the engine computes the minimal set of mutations needed to reach it. It offers file-granularity change tracking, mapping each source file to its own processing component for independent transformation and precise delta detection. The engine natively handles typed multi-dimensional vectors for multimodal AI pipelines and supports elastic distributed indexing that scales to petabyte-scale corpora without manual partitioning.

Cocoindex covers a broad capability surface including building semantic text indexes, constructing knowledge graphs from documents, indexing codebases for AI agents with AST-aware parsing, and serving code context through MCP, CLI, or Claude skills. It can ingest data from any custom source, transform structured and unstructured data together, and export indexed data to local files, cloud storage, or REST APIs. The platform also provides observability tools for tracing data lineage end-to-end and debugging pipeline steps in real time.

The project is configured and extended through Python code, with documentation and installation resources available through its repository.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Features

  • Code Context Servers - Serves a live index of the entire codebase to coding agents through MCP, CLI, or Claude skills.
  • Data Lineage - Tags every output record with source items and transformation version for full provenance tracking.
  • Knowledge Graph Extraction - Parses unstructured text to identify entities, relationships, and statements, storing them as a queryable knowledge graph.
  • Incremental Processing - Ships an incremental processing engine that re-computes only changed data for sub-second freshness.
  • Delta Processing - Re-processes only changed source records or code paths, avoiding full recomputation for sub-second freshness.
  • Python-Defined Transformations - Defines data transformations as pure Python functions and automatically derives the execution graph.
  • Target State Declarations - Implements declarative target-state reconciliation where users describe the desired end state and the engine computes minimal mutations.
  • Declarative Data Reconciliation Pipelines - Provides declarative pipelines that reconcile data stores to a desired state with minimal mutations.
  • Incremental Computation - Processes data changes incrementally so only modified content is re-computed, keeping large corpora fresh without full recomputation.
  • Codebase Indexing - Provides incremental codebase indexing with file-granularity change tracking and AST-aware parsing for AI agents.
  • Knowledge Graph Construction Tools - Builds structured knowledge graphs from documents for AI agent reasoning.
  • Real-Time - Builds real-time knowledge graphs with incremental updates and high-performance graph queries from documents.
  • Multi-Modal RAG Pipelines - Provides a complete pipeline for building multi-modal RAG systems with lineage tracking.
  • Builders - Provides a declarative engine for building semantic text indexes with embeddings and natural language querying.
  • Continuous Sync Engines - Provides a continuous sync engine that watches source changes and keeps derived indexes fresh.
  • Code Search - Provides natural-language code search with syntax-aware chunking for AI agent context retrieval.
  • Schema Inference - Infers output schemas automatically from transformation pipeline type annotations and function signatures.
  • Incremental Vector Sync - Keeps vector indexes continuously updated by processing only the delta from live sources.
  • Codebase Indexing - Maintains a live, shared index of source code that updates automatically with each commit for team agents.
  • AI Agent Indexes - Provides incremental codebase indexing with AST-aware parsing for AI agent context.
  • AST-Aware Parsers - Provides AST-aware parsing to index codebases for AI agent context.
  • Shared Code Index Daemons - Runs a persistent daemon that builds a single code index and serves it to every teammate and agent.
  • Delta-Only Pipeline Executors - Runs only the delta of changed data and code, skipping unchanged work at any scale.
  • Granular File Change Tracking - Maps each source file to its own processing component for independent transformation and precise delta detection.
  • Incremental Data Pipelines - Ships an incremental data pipeline engine that processes only changed data for sub-second freshness at any scale.
  • Data Store Reconciliation - Computes minimal mutations to reconcile a declared target data store state with the current state.
  • Execution Graphs - Composes user-defined Python functions into a directed acyclic graph for scheduled execution with automatic parallelism.
  • Data Provenance Recorders - Records the provenance of every output byte, tracing results back to source records and transformation steps.
  • Automatic ML Workload Batching - Automatically batches GPU and ML workloads like text embeddings for higher throughput.
  • Multi-Vector Embeddings - Natively handles typed multi-dimensional vectors, from simple arrays to multi-vector embeddings for multimodal AI pipelines.
  • Event-Driven Transform Pipelines - Builds real-time transformation pipelines triggered by events from object storage and message queues.
  • Incremental File Converters - Converts source files to target formats incrementally, reprocessing only changed or new files.
  • Data Consistency Models - Implements data consistency models to ensure indexes converge correctly during concurrent updates.
  • Data Ingestion Sources - Reads data from any external system and keeps it incrementally fresh as knowledge for AI agents.
  • Data Destination Connectors - Exports indexed data to any destination including local files, cloud storage, or REST APIs.
  • Multi-Vector Embedding Handlers - Natively handles typed multi-dimensional vectors from simple arrays to multi-vector embeddings for multimodal AI pipelines.
  • File-Level Pipeline Definitions - Maps each source file to its own processing component for independent transformation and precise delta detection.
  • Self-Updating Documentation Generators - Generates and incrementally updates wiki pages for each project in a codebase.
  • Index Scaling - Distributes indexing work across nodes with elastic, fault-tolerant scaling for petabyte-scale corpora.
  • Elastic - Distributes indexing work across nodes with fault-tolerant scaling for petabyte-scale corpora.
  • Unified Structured-Unstructured Indexes - Builds a unified, incrementally updated search index over both structured and unstructured data.
  • Enterprise Source Indexing - Ingests and indexes content from codebases, meetings, Slack, docs, and tickets for live AI agent context.
  • Branch Delta Indexing - Re-indexes only files that differ between branches, layering changes on a shared main index to save compute.
  • Concurrency Controllers - Controls concurrency of data-processing operations to optimize performance and prevent system overload.
  • Cross-Repository Indexes - Provides a daemon that indexes multiple repositories for cross-repo code queries.
  • Data Pipeline Lineage Inspectors - Provides a tool to inspect, trace, and debug every step of a data pipeline in real time.
  • Search Result Provenance Tracers - Traces search results back to source data to debug and improve the indexing strategy.
  • Large File Processing - Processes large files during data indexing by managing granularity, fan-in, fan-out, and memory pressure.
  • RAG Frameworks - ETL framework for indexing data with real-time incremental updates.
  • Data Pipelines - High-performance data transformation framework for AI workflows.
  • Stream Processing - ETL framework designed for building fresh indices for AI applications.
  • Streaming Engines - ETL framework for building real-time AI indexes.
  • Workflow Frameworks - ETL framework for building fresh data indexes.
  • General Productivity Tools - AI-powered search and indexing for codebases.
6,117 stars·449 forks·Rust·apache-2.0·38 views

Star history

Star history chart for cocoindex-io/cocoindexStar history chart for cocoindex-io/cocoindex

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

Frequently asked questions

What does cocoindex-io/cocoindex do?

Cocoindex is an incremental data processing engine that builds and maintains live indexes for AI agents, with a core focus on codebase indexing and knowledge graph extraction. The engine uses a function-graph execution model where user-defined Python functions are composed into a directed acyclic graph, and it processes data incrementally so only changed source records or code paths are re-computed, avoiding full recomputation at any scale. It supports automatic schema…

What are the main features of cocoindex-io/cocoindex?

The main features of cocoindex-io/cocoindex are: Code Context Servers, Data Lineage, Knowledge Graph Extraction, Incremental Processing, Delta Processing, Python-Defined Transformations, Target State Declarations, Declarative Data Reconciliation Pipelines.

Which projects share features with cocoindex-io/cocoindex?

Projects with overlapping indexed features include: cocoindex-io/cocoindex-code — Cocoindex is a command-line code search engine and indexing tool that combines abstract syntax tree pattern matching… zilliztech/claude-context — Claude-context is a retrieval-augmented generation pipeline and semantic code search tool. It functions as an LLM… yifanfeng97/hyper-extract — Hyper-Extract is a framework designed for automated knowledge extraction, graph construction, and retrieval-augmented… zjunlp/deepke — DeepKE is a knowledge extraction toolkit and framework designed to transform unstructured text into structured… eventual-inc/daft — Daft is a distributed dataframe library and multimodal data processor designed to handle large-scale structured and… qq547276542/agriculture_knowledgegraph — Agriculture Knowledge Graph is a structured triple-store system and decision support platform designed to transform…

Projects sharing features with Cocoindex

These projects share indexed features with Cocoindex. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • cocoindex-io/cocoindex-codecocoindex-io avatar

    cocoindex-io/cocoindex-code

    1,962View on GitHub↗

    Cocoindex is a command-line code search engine and indexing tool that combines abstract syntax tree pattern matching and semantic vector embeddings for precise code discovery. It functions locally and integrates with AI coding assistants to automatically retrieve necessary codebase context through standardized communication protocols and persistent background daemon services. The platform employs an asymmetric embedding architecture that generates vector representations using separate parameters for document indexing and search queries. It supports incremental index maintenance by tracking fi

    Pythonagentsastcocoindex
    View on GitHub↗1,962
  • zilliztech/claude-contextzilliztech avatar

    zilliztech/claude-context

    5,373View on GitHub↗

    Claude-context is a retrieval-augmented generation pipeline and semantic code search tool. It functions as an LLM codebase indexer and RAG context provider, designed to index local directories and retrieve relevant code files to provide context for large language models. The system operates as a hybrid search engine that combines keyword matching with dense vector search. This allows for the retrieval of code snippets and logic using natural language queries based on meaning rather than exact text matches. The project covers codebase indexing and search index management, utilizing asynchrono

    TypeScriptagentagentic-ragai-coding
    View on GitHub↗5,373
  • yifanfeng97/hyper-extractyifanfeng97 avatar

    yifanfeng97/Hyper-Extract

    1,242View on GitHub↗

    Hyper-Extract is a framework designed for automated knowledge extraction, graph construction, and retrieval-augmented generation. It functions as a command-line tool that transforms unstructured text into structured knowledge graphs and hypergraphs, enabling users to build interconnected, searchable, and machine-readable data repositories from their documents. The system distinguishes itself through its focus on personal knowledge management and incremental processing. It allows users to update existing knowledge bases by processing only new document deltas, avoiding redundant computation. Th

    Pythonaiai-agentscli
    View on GitHub↗1,242
  • zjunlp/deepkezjunlp avatar

    zjunlp/DeepKE

    4,433View on GitHub↗

    DeepKE is a knowledge extraction toolkit and framework designed to transform unstructured text into structured knowledge graphs. It provides a pipeline for identifying and classifying named entities, semantic relations, and events, converting raw datasets into structured triples. The project utilizes large language models as tool callers through a standardized context protocol to drive automated data extraction processes. It supports schema-driven extraction across multiple domains and bilingual text, employing joint entity and relation extraction to identify components in a single structured

    Python
    View on GitHub↗4,433
Compare all 30 related projects→