awesome-repositories.com
المدونة
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Cinnamon avatar

Cinnamon/kotaemon

0
View on GitHub↗
25,139 نجوم·2,096 تفرعات·Python·apache-2.0·14 مشاهداتcinnamon.github.io/kotaemon↗

Kotaemon

Kotaemon is an orchestration framework designed for building modular, agentic workflows that integrate document processing, retrieval-augmented generation, and multi-step reasoning. It provides a comprehensive platform for developing document-based question answering systems, allowing users to chain language models, prompt templates, and external tools into complex, automated pipelines.

The system distinguishes itself through a highly modular architecture that emphasizes component-based composition and schema-driven data exchange. It supports autonomous agents capable of decomposing complex queries through iterative processing and tool-calling, while its hybrid retrieval orchestration combines vector similarity and full-text search with re-ranking to improve the accuracy of retrieved context. The framework also features event-driven streaming, which delivers incremental results from long-running pipelines to the user interface in real-time.

Beyond its core reasoning capabilities, the platform includes a suite of functional modules for the entire lifecycle of document-based applications. This includes multi-modal parsing for extracting text, tables, and visual elements from diverse file formats, as well as administrative tools for managing document collections, vector stores, and multi-user access. The system is designed to be interface-agnostic, allowing developers to wrap third-party libraries and external services into standardized, reusable processing units.

The project provides a web-based user interface for interactive querying and configuration, and it supports deployment of private, isolated instances through predefined templates.

Features

  • Agentic Reasoning Frameworks - Provides an agentic framework for decomposing complex queries through iterative reasoning and tool-calling.
  • LLM Application Orchestration - Chains language models, prompt templates, and external tools into complex, multi-step reasoning and data processing pipelines.
  • Grounded Answer Generation - Generates grounded answers supported by direct citations and source verification from retrieved documents.
  • Question Answering Systems - Provides an interactive system for retrieving and citing specific evidence from documents to generate grounded answers.
  • Reasoning Chains - Sequences multiple reasoning units into complex workflows with conditional logic for multi-step problem solving.
  • Reasoning Pipelines - Combine prompt templates, language models, and post-processing functions to transform input data into structured outputs.
  • Retrieval-Augmented Generation Frameworks - Constructs modular workflows that combine document indexing, hybrid search, and language models to generate context-aware responses.
  • Conversational Retrieval - Enables interactive chat sessions that retrieve and cite specific evidence from documents to provide grounded answers.
  • Autonomous Agents - Constructs autonomous agents by configuring language models, prompt templates, and tool sets.
  • Agentic Reasoning Frameworks - Implements agentic reasoning loops and tool-use logic to solve complex, multi-hop queries.
  • Embedding Generators - Generates vector representations of text using local or remote models to enable semantic search.
  • Hybrid Search Systems - Combines vector and full-text search with re-ranking to retrieve relevant information for question answering.
  • Document Question Answering Pipelines - Extracts text, tables, and figures from multi-modal documents to support complex, grounded question answering.
  • RAG Frameworks - Provides a modular platform for building document-based question answering systems using LLMs and custom retrieval workflows.
  • Retrieval Orchestration - Orchestrates hybrid retrieval strategies to combine vector and keyword search for improved context accuracy.
  • Language Model Integrations - Connects to various language model providers and local runtimes for document-based question answering.
  • Chain of Thought Implementations - Executes sequential prompt-based steps to break down complex reasoning tasks into manageable parts.
  • Chat Model Integrations - Integrates various chat models through unified interfaces to enable conversational document-based question answering.
  • Tool Calling - Enables autonomous agents to select and execute external tools during conversational interactions.
  • Document Collections - Provides centralized management for document collections to support efficient data organization and retrieval.
  • Semantic Parsing Tools - Extracts text, tables, and visual elements from complex documents using multi-modal parsing.
  • Modular Pipeline Orchestrators - Structures complex data workflows into independent, modular components with caching and logging.
  • Large Language Model Configurations - Provides configuration settings to connect and manage local or cloud-based language and embedding models.
  • Private Document Retrieval - Indexes and queries diverse file formats to provide grounded, cited answers from private document collections.
  • Local Language Model Execution - Loads and executes large language models locally on hardware to enable offline document processing and reasoning.
  • Model Provider Configurations - Centralizes the configuration and authentication of external language and embedding model providers for document processing tasks.
  • Reasoning Workflows - Deploys specialized agent architectures to execute sequential or distributive reasoning steps.
  • Retrieval Strategies - Combines vector similarity, full-text search, and re-ranking to improve retrieval accuracy.
  • Sequential Orchestration - Chains prompts, models, and post-processors into sequential pipelines for automated document processing.
  • Document Parsing Pipelines - Supports multi-modal parsing of diverse file formats to extract text, tables, and visual elements.
  • Vector Document Indexing - Automates the indexing of documents into vector databases to support real-time search and retrieval.
  • Vector Memory Stores - Performs similarity searches against stored embeddings to retrieve relevant context for agentic reasoning.
  • Build Pipeline Extensions - Allows building modular processing workflows by chaining reusable components.
  • Component Composition Patterns - Enables modular pipeline composition by linking reusable processing units into flexible workflows.
  • Web Chat Interfaces - Provides a web-based interface for interactive document querying and conversational AI interactions.
  • Model Provider Integrations - Provides a unified interface to connect local or cloud-based model services for document analysis and retrieval.
  • Citation Management Systems - Provides detailed source references and highlights to verify the accuracy of generated answers.
  • Document Indexing - Retrieves relevant document context from indexed data for use in analytical tasks.
  • External Model Connectors - Configures connections to external language model endpoints to generate text responses for document-based pipelines.
  • Document Chunking Strategies - Segments large documents into manageable chunks to optimize retrieval accuracy for question answering.
  • Language Model Response Generators - Processes chat history through language models to generate conversational responses based on provided context.
  • LLM Application Platforms - Provides an integrated platform for building, testing, and deploying AI-powered applications and workflows.
  • Prompt Templates - Defines reusable prompt structures with dynamic placeholders for language model inputs.
  • Knowledge Retrieval - Clean and customizable UI for document-based chatting.
  • RAG and Data Pipelines - RAG-based tool for chatting with local documents.
  • Retrieval Augmented Generation - Clean, customizable UI for document-based chat.
  • RAG Frameworks - Open-source document Q&A tool with multi-modal support and agentic reasoning.
  • Collaborative Chat Sessions - Supports collaborative chat sessions and shared document access for multiple users.
  • Document Parsing Services - Extracts text and structured content from various file formats using cloud-based analysis services.
  • Data Ingestion - Ingests and parses diverse file formats into structured text for downstream processing and AI consumption.
  • Document Retrieval Interfaces - Provides interfaces for querying document and vector stores to retrieve relevant information for question answering.
  • Vector Database Integrations - Integrates with existing vector database implementations to perform document indexing and similarity searching.
  • Document Stores - Configures storage backends for managing full-text and vector-based document indices.
  • Embedding Service Integrations - Connects to remote embedding services to generate numerical representations of documents for semantic search.
  • Full Text Search - Executes keyword-based full-text search against stored document collections.
  • Search and Indexing - Integrates document storage with full-text search capabilities to enable efficient information discovery.
  • Vector Embedding Indexes - Persists document embeddings in vector databases to enable efficient semantic similarity search.
  • External Tool Integrations - Defines custom tools with input validation to extend agent capabilities with external data sources.
  • Conversation State Managers - Tracks interaction history and manages conversational context to ensure stateful, multi-turn dialogue.
  • Document Rerankers - Filters retrieved content using language models to retain only the most pertinent information.
  • Evidence Extraction Tools - Isolates and retrieves specific document segments that support generated answers to ensure verifiable citations.
  • Inference Endpoint Integrations - Connects to compatible API endpoints to utilize custom or hosted language models for document processing.
  • Local Embedding Generators - Generates vector embeddings locally to enable semantic search without relying on external API dependencies.
  • Model Integration Configurations - Manages programmatic access to configured language and embedding models within custom pipelines.
  • Document Layout Analysis - Parses complex documents to extract text, tables, and visual elements for deep analysis.
  • Vector Databases - Manages the storage and querying of high-dimensional vector embeddings to enable efficient semantic search.
  • Vector Embeddings - Converts text into numerical vector representations using cloud-based models to enable semantic search and document retrieval.
  • Office Document Parsers - Parses common office document formats like PDF, Word, and Excel into structured text nodes for indexing.
  • Web Content Scrapers - Parses live web information into structured formats for use as external context in pipelines.
  • Document and Unstructured Extraction - Extracts text content from various unstructured file formats including office documents, images, and emails.
  • Indexing and Search - Allows customization of indexing and search logic through extensible base classes.
  • Real-Time Data Streaming - Streams incremental results from reasoning pipelines to the user interface in real-time.
  • Vector Collection Management - Performs administrative operations to add, delete, or remove entire collections of embeddings from storage.
  • Tool Wrappers - Exposes internal system components as executable tools for autonomous agents.
  • Web Search Integrations - Fetches real-time information from the internet to provide up-to-date context for question answering.
  • Private Data Hosting - Supports deployment of private, isolated application instances using predefined templates.
  • Retrieval Configuration Interfaces - Provides interfaces to adjust retrieval and generation settings for customized system behavior.
  • Event-Driven Architectures - Implements event-driven processing to deliver incremental pipeline results in real-time.
  • Indexing Pipeline Frameworks - Provides frameworks for extending indexing and reasoning pipelines to meet specific requirements.
  • Custom Component Extensions - Enables the creation of modular indexing and reasoning components for dynamic loading.
  • LLM Response Streaming - Streams generated model output in real-time chunks to provide immediate feedback during long-running interactions.
  • Conditional Execution Flows - Routes pipeline execution based on dynamic evaluation of conditions to handle branching logic.
  • Context-Aware Retrieval - Consolidates retrieved document chunks and media into structured context strings for language models.
  • External Service Integrations - Integrates third-party indexing and transformation tools into the native pipeline architecture.
  • Knowledge Retrieval Tools - Enables agents to retrieve external information by querying the Wikipedia API.
  • Model Credential Managers - Manages administrative credentials for external data sources and language model providers.
  • Document Transformation Pipelines - Provides pipelines for programmatically splitting, filtering, and enriching document collections.
  • Local Document Ingestion - Imports files from local directories with support for recursive scanning and specialized extraction logic.
  • Schema-Driven Data Normalizers - Standardizes heterogeneous data sources into consistent structures to ensure schema uniformity across indexing and reasoning components.
  • Document Retrieval Strategies - Sorts and filters retrieved documents to improve the accuracy of information retrieval.
  • File Upload Management - Handles the upload and storage of user-provided files for subsequent indexing and question-answering interactions.
  • Local Data Persistence - Persists document collections and indexes to local storage to ensure data availability across system restarts.
  • PDF Parsers - Parses PDF files into structured text with spatial coordinates and page metadata.
  • Result Streaming APIs - Streams incremental processing results to provide immediate feedback during long-running tasks.
  • Relevance Ranking Engines - Ranks retrieved documents numerically using language models to improve the relevance of context.
  • Search Result Filtering - Filters search results using language models to remove irrelevant documents before generation.
  • Runtime Configuration Interfaces - Renders user-defined settings in the interface for runtime adjustment of pipeline parameters.
  • Web Content Ingestion Tools - Converts HTML and MHTML web content into structured document objects for processing.
  • Infrastructure Configuration - Defines operational parameters for managing the deployment environment and infrastructure.
  • Pipeline Setting Exposers - Generates user interface controls automatically from custom pipeline configuration parameters.
  • Response Streaming Interfaces - Delivers incremental text responses from language models to the user interface in real-time.

سجل النجوم

مخطط تاريخ النجوم لـ cinnamon/kotaemonمخطط تاريخ النجوم لـ cinnamon/kotaemon

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

الأسئلة الشائعة

ما هي وظيفة cinnamon/kotaemon؟

Kotaemon is an orchestration framework designed for building modular, agentic workflows that integrate document processing, retrieval-augmented generation, and multi-step reasoning. It provides a comprehensive platform for developing document-based question answering systems, allowing users to chain language models, prompt templates, and external tools into complex, automated pipelines.

ما هي الميزات الرئيسية لـ cinnamon/kotaemon؟

الميزات الرئيسية لـ cinnamon/kotaemon هي: Agentic Reasoning Frameworks, LLM Application Orchestration, Grounded Answer Generation, Question Answering Systems, Reasoning Chains, Reasoning Pipelines, Retrieval-Augmented Generation Frameworks, Conversational Retrieval.

ما هي البدائل مفتوحة المصدر لـ cinnamon/kotaemon؟

تشمل البدائل مفتوحة المصدر لـ cinnamon/kotaemon: camel-ai/camel — This project is a comprehensive framework for building and managing autonomous agent systems. It provides a unified… infiniflow/ragflow — This project is a comprehensive retrieval-augmented generation platform designed for building, managing, and deploying… mastra-ai/mastra — Mastra is an orchestration framework designed for building, deploying, and managing autonomous AI agents and… cloudwego/eino — Eino is an AI agent development kit and LLM application framework designed for building autonomous agents and… the-pocket/pocketflow — PocketFlow is a graph-based framework for designing and executing large language model operations and reasoning… vercel/ai — This project is a comprehensive framework for building AI-powered applications, providing a unified toolkit for…

بدائل مفتوحة المصدر لـ Kotaemon

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Kotaemon.
  • camel-ai/camelالصورة الرمزية لـ camel-ai

    camel-ai/camel

    17,253عرض على GitHub↗

    This project is a comprehensive framework for building and managing autonomous agent systems. It provides a unified architecture for orchestrating multi-agent societies, where specialized agents collaborate through roleplay to decompose and solve complex tasks. The system integrates language models with external environments, enabling agents to perform real-world actions through a standardized tool-calling abstraction layer. The framework distinguishes itself through its focus on iterative reasoning and data reliability. It employs automated feedback loops to refine agent outputs and self-eva

    Pythonagentai-societiesartificial-intelligence
    عرض على GitHub↗17,253
  • infiniflow/ragflowالصورة الرمزية لـ infiniflow

    infiniflow/ragflow

    82,922عرض على GitHub↗

    This project is a comprehensive retrieval-augmented generation platform designed for building, managing, and deploying knowledge-based AI applications. It provides a unified environment for organizing datasets, configuring conversational chat assistants, and developing autonomous agents that execute multi-step reasoning workflows. By integrating document intelligence with advanced retrieval pipelines, the platform enables the creation of grounded, verifiable responses supported by traceable citations. The platform distinguishes itself through deep document understanding and sophisticated know

    Pythonagentagenticagentic-ai
    عرض على GitHub↗82,922
mastra-ai/mastraالصورة الرمزية لـ mastra-ai

mastra-ai/mastra

21,221عرض على GitHub↗

Mastra is an orchestration framework designed for building, deploying, and managing autonomous AI agents and multi-agent systems. It provides a comprehensive suite of primitives for creating resilient AI applications, including durable workflow orchestration, event-driven agent loops, and semantic memory management. By integrating these core components, the platform enables developers to build complex, multi-step processes that can reason about goals and execute tasks without manual intervention. The framework distinguishes itself through its focus on observability and secure, isolated execut

TypeScriptagentsaichatbots
عرض على GitHub↗21,221
  • cloudwego/einoالصورة الرمزية لـ cloudwego

    cloudwego/eino

    9,675عرض على GitHub↗

    Eino is an AI agent development kit and LLM application framework designed for building autonomous agents and orchestrating complex language model workflows. It serves as a multi-agent orchestration engine and workflow orchestrator, providing a graph-based execution model to route data between models, tools, and retrievers. The framework distinguishes itself through a robust set of multi-agent coordination patterns, including supervisor-led management, sequential flows, and autonomous reasoning loops like ReAct. It features advanced agent execution controls such as active turn preemption, che

    Goaiai-applicationai-framework
    عرض على GitHub↗9,675
  • عرض جميع البدائل الـ 30 لـ Kotaemon→