awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

34 रिपॉजिटरी

Awesome GitHub RepositoriesKnowledge Retrieval

Methods for integrating external knowledge into the extraction process.

Explore 34 awesome GitHub repositories matching part of an awesome list · Knowledge Retrieval. Refine with filters or upvote what's useful.

Awesome Knowledge Retrieval GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • langgenius/difylanggenius का अवतार

    langgenius/dify

    145,458GitHub पर देखें↗

    Dify is an open-source platform for building, orchestrating, and deploying generative AI applications and autonomous agents. It provides a visual development environment that allows users to design complex, multi-step logic chains and conversational flows, which can then be published as APIs, web interfaces, or embedded widgets. The platform acts as a centralized infrastructure layer, managing model connections, prompt templates, and knowledge retrieval to support scalable AI-powered services. What distinguishes the platform is its focus on stateful application design and workflow orchestrati

    Platform for developing LLM applications with RAG pipelines.

    TypeScriptagentagentic-aiagentic-framework
    GitHub पर देखें↗145,458
  • infiniflow/ragflowinfiniflow का अवतार

    infiniflow/ragflow

    82,922GitHub पर देखें↗

    This project is a comprehensive retrieval-augmented generation platform designed for building, managing, and deploying knowledge-based AI applications. It provides a unified environment for organizing datasets, configuring conversational chat assistants, and developing autonomous agents that execute multi-step reasoning workflows. By integrating document intelligence with advanced retrieval pipelines, the platform enables the creation of grounded, verifiable responses supported by traceable citations. The platform distinguishes itself through deep document understanding and sophisticated know

    RAG engine based on deep document understanding.

    Pythonagentagenticagentic-ai
    GitHub पर देखें↗82,922
  • mintplex-labs/anything-llmMintplex-Labs का अवतार

    Mintplex-Labs/anything-llm

    61,663GitHub पर देखें↗

    This platform serves as a comprehensive environment for managing private language models, document knowledge bases, and automated agent workflows within secure local infrastructure. It functions as a document-aware workspace that enables users to ingest diverse file formats into searchable repositories, ensuring that all data processing and model inference remain within private, local environments to maintain data sovereignty. The system distinguishes itself through a modular agentic engine that allows for the definition of custom skills and external tool execution. By utilizing a multi-model

    All-in-one application with full RAG and agent capabilities.

    JavaScriptai-agentscustom-ai-agentsdeepseek
    GitHub पर देखें↗61,663
  • quivrhq/quivrQuivrHQ का अवतार

    QuivrHQ/quivr

    39,165GitHub पर देखें↗

    Quivr is a retrieval-augmented generation platform designed to transform raw documents into searchable knowledge bases. It functions as a centralized environment where users can ingest files, index them into vector databases, and interact with language models to receive contextually relevant, data-backed responses. The platform distinguishes itself through an agentic workflow orchestrator that sequences retrieval tasks, tool execution, and model interactions to resolve complex, multi-step queries. This engine is entirely configuration-driven, allowing users to define document ingestion, chunk

    Personal productivity assistant for chatting with local documents.

    Pythonaiapichatbot
    GitHub पर देखें↗39,165
  • chatchat-space/langchain-chatchatchatchat-space का अवतार

    chatchat-space/Langchain-Chatchat

    38,211GitHub पर देखें↗

    Langchain-Chatchat is a system for building retrieval-augmented generation applications and autonomous AI agents. It integrates a knowledge base management system and an agent framework to enable language models to interact with private documents and execute multi-step tasks through external tools. The platform supports local deployment of language models on private infrastructure to operate without an internet connection. It includes a multimodal AI platform that combines vision models for image analysis with text-to-image generation capabilities. The system provides a web-based conversatio

    Local knowledge base Q&A using various language models.

    Pythonchatbotchatchatchatglm
    GitHub पर देखें↗38,211
  • hkuds/lightragHKUDS का अवतार

    HKUDS/LightRAG

    36,651GitHub पर देखें↗

    LightRAG is a graph-based retrieval framework designed to build retrieval-augmented generation pipelines. It structures unstructured text into knowledge graphs, enabling multi-hop reasoning and complex query synthesis across large document collections. By integrating dense vector embeddings with structured knowledge graphs, the system facilitates both similarity-based and relationship-aware information retrieval. The framework distinguishes itself through a dual-level retrieval strategy that combines low-level keyword matching with high-level semantic graph traversal to capture both specific

    Simple and fast retrieval-augmented generation framework.

    Pythongenaigptgpt-4
    GitHub पर देखें↗36,651
  • microsoft/graphragmicrosoft का अवतार

    microsoft/graphrag

    33,792GitHub पर देखें↗

    GraphRAG is a data processing pipeline and retrieval engine designed to transform unstructured text into interconnected knowledge graphs. By utilizing language models to extract entities and relationships, it builds structured representations of information that enable context-aware retrieval for downstream applications. The system distinguishes itself through hierarchical graph clustering and large-scale data synthesis, which organize massive document corpora into multi-level structures. This approach allows for both vector-based semantic searches and graph-based traversals, providing a comp

    Modular graph-based retrieval-augmented generation system.

    Pythongptgpt-4gpt4
    GitHub पर देखें↗33,792
  • labring/fastgptlabring का अवतार

    labring/FastGPT

    27,132GitHub पर देखें↗

    FastGPT is a comprehensive platform for building, deploying, and managing context-aware artificial intelligence applications. It provides a unified environment that integrates custom data sources with language models, utilizing a retrieval-augmented generation engine to ground responses in accurate, domain-specific information. The system is designed for enterprise-scale use, featuring multi-tenant architecture, administrative controls, and secure authentication protocols including OAuth 2.0 and custom single sign-on integration. The platform distinguishes itself through a visual, node-based

    Knowledge-based platform with visual workflow orchestration.

    TypeScriptagentclaudedeepseek
    GitHub पर देखें↗27,132
  • cinnamon/kotaemonCinnamon का अवतार

    Cinnamon/kotaemon

    25,139GitHub पर देखें↗

    Kotaemon is an orchestration framework designed for building modular, agentic workflows that integrate document processing, retrieval-augmented generation, and multi-step reasoning. It provides a comprehensive platform for developing document-based question answering systems, allowing users to chain language models, prompt templates, and external tools into complex, automated pipelines. The system distinguishes itself through a highly modular architecture that emphasizes component-based composition and schema-driven data exchange. It supports autonomous agents capable of decomposing complex q

    Clean and customizable UI for document-based chatting.

    Pythonchatbotllmsopen-source
    GitHub पर देखें↗25,139
  • 1panel-dev/maxkb1Panel-dev का अवतार

    1Panel-dev/MaxKB

    21,337GitHub पर देखें↗

    MaxKB is a self-hosted retrieval-augmented generation platform designed to connect internal document repositories with large language models. It functions as an enterprise knowledge management system that enables organizations to query private data through a conversational interface, providing automated responses based on uploaded files and internal business information. The platform distinguishes itself by normalizing diverse data sources into a unified index, which is then processed through chunking and vector-based retrieval to ensure context-aware results. It manages session state and pro

    Out-of-the-box knowledge base question answering system.

    Pythonagentagentic-aichatbot
    GitHub पर देखें↗21,337
  • hkuds/rag-anythingHKUDS का अवतार

    HKUDS/RAG-Anything

    21,372GitHub पर देखें↗

    RAG-Anything is a retrieval-augmented generation framework designed to index diverse document formats and perform semantic search using local machine learning models. It functions as a local multimodal data processor, extracting and organizing information from various file types into a unified knowledge base to facilitate private document analysis. The system distinguishes itself through its high-throughput ingestion engine, which processes large batches of documents into searchable vector embeddings. By executing machine learning models directly on local hardware, the framework ensures that

    All-in-one system for retrieval-augmented generation.

    Pythonmulti-modal-ragretrieval-augmented-generation
    GitHub पर देखें↗21,372
  • eosphoros-ai/db-gpteosphoros-ai का अवतार

    eosphoros-ai/DB-GPT

    18,999GitHub पर देखें↗

    DB-GPT is an agentic data analysis platform and business intelligence AI that functions as a large language model data assistant. It provides a text-to-SQL interface and a sandboxed code execution environment to translate natural language into executable database queries and Python scripts. The platform utilizes iterative agentic reasoning to plan and execute multi-step data analysis workflows through tool calls. It features a modular skill-based extension system that allows domain knowledge and analysis workflows to be packaged into reusable functional components. The system integrates data

    GraphRAG integrating knowledge graphs and document structures.

    Pythonagentsbgidatabase
    GitHub पर देखें↗18,999
  • explodinggradients/ragasexplodinggradients का अवतार

    explodinggradients/ragas

    14,400GitHub पर देखें↗

    Ragas is an evaluation framework and performance benchmark designed to quantify the quality of retrieval augmented generation pipelines. It functions as an application optimizer to identify bottlenecks in language model workflows using automated metrics and model-based scoring. The framework includes a system for generating synthetic datasets that mimic production scenarios and edge cases to create realistic test cases. It enables reference-free assessment, allowing the evaluation of response quality by analyzing grounding in the provided context without requiring gold-standard labels. The s

    Evaluation framework for RAG pipeline components.

    Python
    GitHub पर देखें↗14,400
  • netease-youdao/qanythingnetease-youdao का अवतार

    netease-youdao/QAnything

    14,020GitHub पर देखें↗

    QAnything is a retrieval-augmented generation application framework and self-hosted AI interface. It functions as a system that combines a vector database knowledge base, a document parsing service, and a hybrid search engine to generate answers based on private user data. The project features a modular pipeline architecture that allows users to independently replace components such as parsers, embedding models, and reranking engines. It supports local-first model deployment and offline operation to ensure data privacy, and includes a two-stage retrieval pipeline that merges dense vector embe

    Question and answer system for arbitrary document types.

    Python
    GitHub पर देखें↗14,020
  • openspg/kagOpenSPG का अवतार

    OpenSPG/KAG

    8,548GitHub पर देखें↗

    KAG is a graph-augmented retrieval augmented generation system and knowledge graph engine. It functions as a framework that integrates large language models with graph retrieval and numerical calculation to resolve natural language queries. The system creates unified knowledge representations by aligning unstructured data and expert rules through semantic mapping. It maintains mutual indexing between graph structures and original text blocks to ensure that reasoning processes remain linked to verifiable source data. The project provides capabilities for semantic information integration, grap

    Knowledge-enhanced generation framework for rigorous decision-making.

    Pythonknowledge-graphlarge-language-modellogical-reasoning
    GitHub पर देखें↗8,548
  • weaviate/verbaweaviate का अवतार

    weaviate/Verba

    7,715GitHub पर देखें↗

    Verba is a retrieval-augmented generation interface and chatbot that uses Weaviate to provide factual answers based on private datasets. It functions as a vector database knowledge base, combining a hybrid search engine with an orchestration interface to connect various large language model providers and embedding services. The system differentiates itself through a RAG pipeline manager for adjusting text chunking rules and retrieval settings, alongside a 3D vector space visualization tool for analyzing the spatial organization and clustering of high-dimensional embeddings. It employs a modul

    RAG chatbot powered by the Weaviate vector database.

    Python
    GitHub पर देखें↗7,715
  • marker-inc-korea/autoragMarker-Inc-Korea का अवतार

    Marker-Inc-Korea/AutoRAG

    4,833GitHub पर देखें↗

    AutoRAG is an automation layer and optimization tool for retrieval-augmented generation. It provides a framework for measuring pipeline performance through an evaluation system and an automated search strategy that identifies the most effective combinations of retrieval and generation modules. The system distinguishes itself through AutoML-style optimization, using hyperparameter grid searches and automated trials to find the highest performing architectural configuration for a specific dataset. It includes a specialized dataset generator that creates synthetic question-answer pairs and groun

    AutoML tool for finding optimal RAG pipelines.

    Python
    GitHub पर देखें↗4,833
  • ragapp/ragappragapp का अवतार

    ragapp/ragapp

    4,438GitHub पर देखें↗

    This project is an agentic retrieval-augmented generation platform and orchestration framework designed to connect large language models to private enterprise data. It serves as a self-hosted AI gateway that integrates vector databases and external tools to automate complex information retrieval and generation tasks. The system differentiates itself through an AI agent workflow builder that orchestrates multiple specialized agents with distinct roles to solve multi-step problems. It includes a dedicated vector database integration interface for indexing private documents and a secure sandbox

    Enterprise-ready framework for agentic RAG workflows.

    TypeScript
    GitHub पर देखें↗4,438
  • gusye1234/nano-graphraggusye1234 का अवतार

    gusye1234/nano-graphrag

    3,896GitHub पर देखें↗

    nano-graphrag एक रिट्रीवल सिस्टम है जो लार्ज लैंग्वेज मॉडल प्रतिक्रियाओं के लिए संरचित संदर्भ प्रदान करने के लिए नॉलेज ग्राफ का उपयोग करता है। यह एक नॉलेज ग्राफ इंडेक्सर के रूप में कार्य करता है जो असंरचित टेक्स्ट को संस्थाओं और संबंधों के नेटवर्क में बदलता है, साथ ही एक हाइब्रिड ग्राफ रिट्रीवल सिस्टम के रूप में भी कार्य करता है। यह प्रोजेक्ट जटिल प्राकृतिक भाषा के सवालों के जवाब देने के लिए स्थानीय पड़ोस खोजों को वैश्विक सामुदायिक सारांशों के साथ जोड़कर खुद को अलग करता है। इसमें एक नॉलेज ग्राफ विज़ुअलाइज़र शामिल है जो अनुक्रमित ज्ञान को मैप करने के लिए संस्थाओं और उनके संबंधों के HTML प्रतिनिधित्व उत्पन्न करता है। यह फ्रेमवर्क संस्था-संबंध निष्कर्षण, समुदाय-आधारित ग्राफ क्लस्टरिंग और हैश-आधारित इंक्रीमेंटल इंडेक्सिंग सहित क्षमताओं के एक व्यापक सेट को कवर करता है। यह ओपन-सोर्स मॉडल और स्थानीय एम्बेडिंग प्रदाताओं को जोड़ने के लिए एक एकीकरण परत प्रदान करता है, जो की-वैल्यू, वेक्टर और ग्राफ डेटा के लिए प्लगेबल स्टोरेज बैकएंड द्वारा समर्थित है। अतिरिक्त उपयोगिता तर्क-आधारित प्रतिक्रिया कैशिंग और भाषा मॉडल से अस्थिर JSON आउटपुट की मरम्मत के लिए पोस्ट-प्रोसेसिंग कार्यों के माध्यम से प्रदान की जाती है।

    Simple and hackable implementation of GraphRAG.

    Python
    GitHub पर देखें↗3,896
  • circlemind-ai/fast-graphragcirclemind-ai का अवतार

    circlemind-ai/fast-graphrag

    3,811GitHub पर देखें↗

    Fast-GraphRAG is a system for generating and querying knowledge graphs from domain data. It uses a GraphRAG retrieval workflow to traverse structured data and isolate precise evidence for answering complex questions. The project utilizes an agent-driven retrieval framework to coordinate the querying of knowledge graphs and the synthesis of final answers. It supports incremental data synchronization, allowing structured knowledge bases to be updated in real time as source information evolves. The system integrates with API-compatible language models and embedding providers to power its data p

    Adaptive RAG that adjusts to specific data and queries.

    Python
    GitHub पर देखें↗3,811
पिछला12अगला
  1. Home
  2. Part of an Awesome List
  3. AI & Machine Learning
  4. Knowledge Retrieval