awesome-repositories.com
Blog
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
cocoindex-io avatar

cocoindex-io/cocoindex

0
View on GitHub↗
cocoindex.io↗

Cocoindex

Cocoindex is an incremental data processing engine that builds and maintains live indexes for AI agents, with a core focus on codebase indexing and knowledge graph extraction. The engine uses a function-graph execution model where user-defined Python functions are composed into a directed acyclic graph, and it processes data incrementally so only changed source records or code paths are re-computed, avoiding full recomputation at any scale. It supports automatic schema inference from transformation pipeline type annotations and provides full data lineage tracing, tagging every output record with its source items and transformation version.

The project distinguishes itself through declarative target-state reconciliation, where users describe the desired end state of a data store in Python and the engine computes the minimal set of mutations needed to reach it. It offers file-granularity change tracking, mapping each source file to its own processing component for independent transformation and precise delta detection. The engine natively handles typed multi-dimensional vectors for multimodal AI pipelines and supports elastic distributed indexing that scales to petabyte-scale corpora without manual partitioning.

Cocoindex covers a broad capability surface including building semantic text indexes, constructing knowledge graphs from documents, indexing codebases for AI agents with AST-aware parsing, and serving code context through MCP, CLI, or Claude skills. It can ingest data from any custom source, transform structured and unstructured data together, and export indexed data to local files, cloud storage, or REST APIs. The platform also provides observability tools for tracing data lineage end-to-end and debugging pipeline steps in real time.

The project is configured and extended through Python code, with documentation and installation resources available through its repository.

Features

  • Code Context Servers - Serves a live index of the entire codebase to coding agents through MCP, CLI, or Claude skills.
  • Data Lineage - Tags every output record with source items and transformation version for full provenance tracking.
  • Knowledge Graph Extraction - Parses unstructured text to identify entities, relationships, and statements, storing them as a queryable knowledge graph.
  • Incremental Processing - Ships an incremental processing engine that re-computes only changed data for sub-second freshness.

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI
6,117 Stars·449 Forks·Rust·apache-2.0·10 Aufrufe
  • Delta Processing - Re-processes only changed source records or code paths, avoiding full recomputation for sub-second freshness.
  • Python-Defined Transformations - Defines data transformations as pure Python functions and automatically derives the execution graph.
  • Target State Declarations - Implements declarative target-state reconciliation where users describe the desired end state and the engine computes minimal mutations.
  • Declarative Data Reconciliation Pipelines - Provides declarative pipelines that reconcile data stores to a desired state with minimal mutations.
  • Incremental Computation - Processes data changes incrementally so only modified content is re-computed, keeping large corpora fresh without full recomputation.
  • Codebase Indexing - Provides incremental codebase indexing with file-granularity change tracking and AST-aware parsing for AI agents.
  • Knowledge Graph Construction Tools - Builds structured knowledge graphs from documents for AI agent reasoning.
  • Real-Time - Builds real-time knowledge graphs with incremental updates and high-performance graph queries from documents.
  • Multi-Modal RAG Pipelines - Provides a complete pipeline for building multi-modal RAG systems with lineage tracking.
  • Builders - Provides a declarative engine for building semantic text indexes with embeddings and natural language querying.
  • Continuous Sync Engines - Provides a continuous sync engine that watches source changes and keeps derived indexes fresh.
  • Code Search - Provides natural-language code search with syntax-aware chunking for AI agent context retrieval.
  • Schema Inference - Infers output schemas automatically from transformation pipeline type annotations and function signatures.
  • Incremental Vector Sync - Keeps vector indexes continuously updated by processing only the delta from live sources.
  • Codebase Indexing - Maintains a live, shared index of source code that updates automatically with each commit for team agents.
  • AI Agent Indexes - Provides incremental codebase indexing with AST-aware parsing for AI agent context.
  • AST-Aware Parsers - Provides AST-aware parsing to index codebases for AI agent context.
  • Shared Code Index Daemons - Runs a persistent daemon that builds a single code index and serves it to every teammate and agent.
  • Delta-Only Pipeline Executors - Runs only the delta of changed data and code, skipping unchanged work at any scale.
  • Granular File Change Tracking - Maps each source file to its own processing component for independent transformation and precise delta detection.
  • Incremental Data Pipelines - Ships an incremental data pipeline engine that processes only changed data for sub-second freshness at any scale.
  • Data Store Reconciliation - Computes minimal mutations to reconcile a declared target data store state with the current state.
  • Execution Graphs - Composes user-defined Python functions into a directed acyclic graph for scheduled execution with automatic parallelism.
  • Data Provenance Recorders - Records the provenance of every output byte, tracing results back to source records and transformation steps.
  • Automatic ML Workload Batching - Automatically batches GPU and ML workloads like text embeddings for higher throughput.
  • Multi-Vector Embeddings - Natively handles typed multi-dimensional vectors, from simple arrays to multi-vector embeddings for multimodal AI pipelines.
  • Event-Driven Transform Pipelines - Builds real-time transformation pipelines triggered by events from object storage and message queues.
  • Incremental File Converters - Converts source files to target formats incrementally, reprocessing only changed or new files.
  • Data Consistency Models - Implements data consistency models to ensure indexes converge correctly during concurrent updates.
  • Data Ingestion Sources - Reads data from any external system and keeps it incrementally fresh as knowledge for AI agents.
  • Data Destination Connectors - Exports indexed data to any destination including local files, cloud storage, or REST APIs.
  • Multi-Vector Embedding Handlers - Natively handles typed multi-dimensional vectors from simple arrays to multi-vector embeddings for multimodal AI pipelines.
  • File-Level Pipeline Definitions - Maps each source file to its own processing component for independent transformation and precise delta detection.
  • Self-Updating Documentation Generators - Generates and incrementally updates wiki pages for each project in a codebase.
  • Index Scaling - Distributes indexing work across nodes with elastic, fault-tolerant scaling for petabyte-scale corpora.
  • Elastic - Distributes indexing work across nodes with fault-tolerant scaling for petabyte-scale corpora.
  • Unified Structured-Unstructured Indexes - Builds a unified, incrementally updated search index over both structured and unstructured data.
  • Enterprise Source Indexing - Ingests and indexes content from codebases, meetings, Slack, docs, and tickets for live AI agent context.
  • Branch Delta Indexing - Re-indexes only files that differ between branches, layering changes on a shared main index to save compute.
  • Concurrency Controllers - Controls concurrency of data-processing operations to optimize performance and prevent system overload.
  • Cross-Repository Indexes - Provides a daemon that indexes multiple repositories for cross-repo code queries.
  • Data Pipeline Lineage Inspectors - Provides a tool to inspect, trace, and debug every step of a data pipeline in real time.
  • Search Result Provenance Tracers - Traces search results back to source data to debug and improve the indexing strategy.
  • Large File Processing - Processes large files during data indexing by managing granularity, fan-in, fan-out, and memory pressure.
  • RAG Frameworks - ETL framework for indexing data with real-time incremental updates.
  • Data Pipelines - High-performance data transformation framework for AI workflows.
  • Stream Processing - ETL framework designed for building fresh indices for AI applications.
  • Streaming Engines - ETL framework for building real-time AI indexes.
  • Workflow Frameworks - ETL framework for building fresh data indexes.
  • General Productivity Tools - AI-powered search and indexing for codebases.
  • Star-Verlauf

    Star-Verlauf für cocoindex-io/cocoindexStar-Verlauf für cocoindex-io/cocoindex

    Open-Source-Alternativen zu Cocoindex

    Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Cocoindex.
    • zilliztech/claude-contextAvatar von zilliztech

      zilliztech/claude-context

      5,373Auf GitHub ansehen↗

      Claude-context is a retrieval-augmented generation pipeline and semantic code search tool. It functions as an LLM codebase indexer and RAG context provider, designed to index local directories and retrieve relevant code files to provide context for large language models. The system operates as a hybrid search engine that combines keyword matching with dense vector search. This allows for the retrieval of code snippets and logic using natural language queries based on meaning rather than exact text matches. The project covers codebase indexing and search index management, utilizing asynchrono

      TypeScriptagentagentic-ragai-coding
      Auf GitHub ansehen↗5,373
    • yifanfeng97/hyper-extractAvatar von yifanfeng97

      yifanfeng97/Hyper-Extract

      1,242Auf GitHub ansehen↗

      Hyper-Extract is a framework designed for automated knowledge extraction, graph construction, and retrieval-augmented generation. It functions as a command-line tool that transforms unstructured text into structured knowledge graphs and hypergraphs, enabling users to build interconnected, searchable, and machine-readable data repositories from their documents. The system distinguishes itself through its focus on personal knowledge management and incremental processing. It allows users to update existing knowledge bases by processing only new document deltas, avoiding redundant computation. Th

      Pythonaiai-agentscli
      Auf GitHub ansehen↗1,242
    • zjunlp/deepkeAvatar von zjunlp

      zjunlp/DeepKE

      4,433Auf GitHub ansehen↗

      DeepKE is a knowledge extraction toolkit and framework designed to transform unstructured text into structured knowledge graphs. It provides a pipeline for identifying and classifying named entities, semantic relations, and events, converting raw datasets into structured triples. The project utilizes large language models as tool callers through a standardized context protocol to drive automated data extraction processes. It supports schema-driven extraction across multiple domains and bilingual text, employing joint entity and relation extraction to identify components in a single structured

      Python
      Auf GitHub ansehen↗4,433
    • eventual-inc/daftAvatar von Eventual-Inc

      Eventual-Inc/Daft

      5,225Auf GitHub ansehen↗

      Daft is a distributed dataframe library and multimodal data processor designed to handle large-scale structured and unstructured data. It functions as a vectorized execution engine that processes tables alongside images, audio, and video, utilizing a unified schema to manage diverse data types. The project distinguishes itself by combining distributed data engineering with large-scale AI inference. It provides an AI data pipeline for batch-optimizing model prompts and generating high-dimensional text embeddings, while utilizing zero-copy memory sharing to execute custom Python functions witho

      Rustai-engineeringai-pipelinearrow
      Auf GitHub ansehen↗5,225
    Alle 30 Alternativen zu Cocoindex anzeigen→

    Häufig gestellte Fragen

    Was macht cocoindex-io/cocoindex?

    Cocoindex is an incremental data processing engine that builds and maintains live indexes for AI agents, with a core focus on codebase indexing and knowledge graph extraction. The engine uses a function-graph execution model where user-defined Python functions are composed into a directed acyclic graph, and it processes data incrementally so only changed source records or code paths are re-computed, avoiding full recomputation at any scale. It supports automatic schema…

    Was sind die Hauptfunktionen von cocoindex-io/cocoindex?

    Die Hauptfunktionen von cocoindex-io/cocoindex sind: Code Context Servers, Data Lineage, Knowledge Graph Extraction, Incremental Processing, Delta Processing, Python-Defined Transformations, Target State Declarations, Declarative Data Reconciliation Pipelines.

    Welche Open-Source-Alternativen gibt es zu cocoindex-io/cocoindex?

    Open-Source-Alternativen zu cocoindex-io/cocoindex sind unter anderem: zilliztech/claude-context — Claude-context is a retrieval-augmented generation pipeline and semantic code search tool. It functions as an LLM… yifanfeng97/hyper-extract — Hyper-Extract is a framework designed for automated knowledge extraction, graph construction, and retrieval-augmented… zjunlp/deepke — DeepKE is a knowledge extraction toolkit and framework designed to transform unstructured text into structured… eventual-inc/daft — Daft is a distributed dataframe library and multimodal data processor designed to handle large-scale structured and… qq547276542/agriculture_knowledgegraph — Agriculture Knowledge Graph is a structured triple-store system and decision support platform designed to transform… unionai-oss/pandera — Pandera is a data pipeline validation framework and statistical type validation tool. It functions as a library for…