awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目关于排名机制媒体报道MCP 服务器
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

Enterprise knowledge retrieval

排名更新于 2026年7月15日

For enterprise knowledge retrieval, the strongest matches are onyx-dot-app/onyx (Onyx is a comprehensive enterprise search and retrieval platform), truefoundry/cognita (Cognita is a RAG orchestration framework that provides the) and danswer-ai/danswer (Danswer is a self-hosted enterprise search and RAG platform). hkuds/rag-anything and modsetter/surfsense round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

Find the best enterprise knowledge retrieval tools. We compare top open-source RAG frameworks by activity and features to help you pick the right one.

Enterprise knowledge retrieval

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • onyx-dot-app/onyxonyx-dot-app 的头像

    onyx-dot-app/onyx

    17,491在 GitHub 上查看↗

    Onyx is an enterprise-grade AI platform designed for knowledge management, search, and autonomous agent orchestration. It functions as a centralized system that aggregates unstructured organizational data, enabling secure, context-aware retrieval and interaction across internal documents and communication history. By integrating retrieval-augmented generation with multi-model orchestration, the platform provides a unified interface for teams to query internal knowledge bases and execute complex, multi-step business processes. The platform distinguishes itself through a focus on private infras

    Onyx is a comprehensive enterprise search and retrieval platform that natively supports RAG pipelines, document parsing, multi-source connectors, and role-based access control, all while being fully self-hostable.

    PythonOn-Premise DeploymentRole-Based Access ControlRetrieval Augmented Generation Systems
    在 GitHub 上查看↗17,491
  • truefoundry/cognitatruefoundry 的头像

    truefoundry/cognita

    4,317在 GitHub 上查看↗

    Cognita is a retrieval augmented generation orchestration framework used to build pipelines that connect document stores and language models to provide grounded answers. It functions as a document ingestion pipeline and a vector database integrator, managing the process of loading, parsing, and indexing files into a searchable knowledge base. The system includes a language model gateway proxy that provides a unified API to interact with multiple different model providers. This routing layer decouples the application from specific vendors, allowing requests to be proxied through a provider-agn

    Cognita is a RAG orchestration framework that provides the necessary document parsing, vector database integration, and retrieval pipelines to build an enterprise search system, though it functions as a developer-focused toolkit rather than a pre-packaged, ready-to-deploy application.

    PythonDocument Ingestion PipelinesDocument Ingestion PipelinesRetrieval-Augmented Generation Frameworks
    在 GitHub 上查看↗4,317
  • danswer-ai/danswerdanswer-ai 的头像

    danswer-ai/danswer

    30,552在 GitHub 上查看↗

    Danswer is an LLM application framework and RAG engine that provides a self-hosted interface for connecting large language models to private data. It serves as an enterprise AI chat interface and agent orchestrator, enabling the creation of specialized assistants with custom instructions and knowledge bases. The platform differentiates itself through an observability dashboard for tracking query history and token consumption, as well as a white-labeled interface for customized branding. It includes a multi-step research workflow for producing long-form reports and a sandboxed environment for

    Danswer is a self-hosted enterprise search and RAG platform that natively integrates document parsing, multi-source connectors, and role-based access control to enable secure querying of internal company data.

    PythonData ConnectorsRole-Based Access ControlRole-Based Access Controls
    在 GitHub 上查看↗30,552
  • hkuds/rag-anythingHKUDS 的头像

    HKUDS/RAG-Anything

    21,372在 GitHub 上查看↗

    RAG-Anything is a retrieval-augmented generation framework designed to index diverse document formats and perform semantic search using local machine learning models. It functions as a local multimodal data processor, extracting and organizing information from various file types into a unified knowledge base to facilitate private document analysis. The system distinguishes itself through its high-throughput ingestion engine, which processes large batches of documents into searchable vector embeddings. By executing machine learning models directly on local hardware, the framework ensures that

    This framework provides the core RAG pipeline, document parsing, and vector database integration needed for internal data retrieval, though it lacks explicit mention of enterprise-grade role-based access control and multi-source connectors.

    PythonDocument Parsing PipelinesRetrieval-Augmented Generation FrameworksVector Databases
    在 GitHub 上查看↗21,372
  • modsetter/surfsenseMODSetter 的头像

    MODSetter/SurfSense

    14,816在 GitHub 上查看↗

    SurfSense is a self-hosted platform designed for building retrieval-augmented generation pipelines and managing private knowledge bases. It functions as a containerized research stack that allows users to index diverse data sources and query them using language models, ensuring that all information retrieval is grounded in specific source citations. The platform distinguishes itself through its modular architecture, which supports the integration of custom tools and diverse language models via a unified abstraction layer. It facilitates secure, collaborative research environments by implement

    SurfSense is a self-hosted platform built specifically for RAG pipelines and knowledge base management, providing the core indexing and retrieval capabilities required for an enterprise search system.

    PythonRetrieval Augmented Generation PipelinesRetrieval-Augmented Generation FrameworksRole-Based Access Control
    在 GitHub 上查看↗14,816
  • quivrhq/quivrQuivrHQ 的头像

    QuivrHQ/quivr

    39,165在 GitHub 上查看↗

    Quivr is a retrieval-augmented generation platform designed to transform raw documents into searchable knowledge bases. It functions as a centralized environment where users can ingest files, index them into vector databases, and interact with language models to receive contextually relevant, data-backed responses. The platform distinguishes itself through an agentic workflow orchestrator that sequences retrieval tasks, tool execution, and model interactions to resolve complex, multi-step queries. This engine is entirely configuration-driven, allowing users to define document ingestion, chunk

    Quivr is a self-hostable RAG platform that provides document ingestion, vector indexing, and an agentic interface for querying internal knowledge bases, making it a direct fit for an enterprise retrieval system.

    PythonDocument Ingestion PipelinesRetrieval-Augmented Generation FrameworksRetrieval Augmented Generation Systems
    在 GitHub 上查看↗39,165
  • mintplex-labs/anything-llmMintplex-Labs 的头像

    Mintplex-Labs/anything-llm

    61,663在 GitHub 上查看↗

    This platform serves as a comprehensive environment for managing private language models, document knowledge bases, and automated agent workflows within secure local infrastructure. It functions as a document-aware workspace that enables users to ingest diverse file formats into searchable repositories, ensuring that all data processing and model inference remain within private, local environments to maintain data sovereignty. The system distinguishes itself through a modular agentic engine that allows for the definition of custom skills and external tool execution. By utilizing a multi-model

    This platform provides a self-hostable, document-aware workspace that integrates vector databases, RAG pipelines, and document parsing to enable secure, private retrieval over internal data.

    JavaScriptDocument Ingestion PipelinesDocument Parsing PipelinesRetrieval Augmented Generation Systems
    在 GitHub 上查看↗61,663
  • zylon-ai/private-gptzylon-ai 的头像

    zylon-ai/private-gpt

    57,278在 GitHub 上查看↗

    This project is a privacy-first backend service designed to facilitate retrieval-augmented generation by processing local documents into searchable vector representations. It provides a modular architecture that allows users to ingest diverse file formats, manage document metadata, and perform semantic searches to provide context-aware responses for chat and completion requests. The system distinguishes itself through a database-agnostic abstraction layer that supports various storage backends, ranging from local disk storage to enterprise-grade vector databases. It offers flexible deployment

    This project provides a privacy-focused RAG pipeline and document ingestion system that functions as a core backend for enterprise search and retrieval, though it lacks built-in role-based access control.

    PythonDocument Ingestion PipelinesDocument Parsing PipelinesRetrieval-Augmented Generation Pipelines
    在 GitHub 上查看↗57,278
  • asyncfuncai/deepwiki-openAsyncFuncAI 的头像

    AsyncFuncAI/deepwiki-open

    14,362在 GitHub 上查看↗

    This platform is an automated documentation and codebase analysis system designed to generate structured wikis, technical guides, and interactive diagrams from source code repositories. It functions as a retrieval-augmented generation framework that connects codebases to language models, enabling context-aware answers, deep research, and automated documentation updates through semantic vector search. The system distinguishes itself through a self-hosted, containerized architecture that supports both cloud-based and local AI model execution. It provides sophisticated model orchestration, allow

    This system functions as a self-hosted RAG pipeline specifically designed to index and query technical documentation and codebases, making it a direct fit for enterprise retrieval needs despite its specialized focus on developer-centric content.

    PythonRetrieval Augmented GenerationRetrieval Augmented Generation PipelinesRetrieval-Augmented Generation Frameworks
    在 GitHub 上查看↗14,362
  • infiniflow/ragflowinfiniflow 的头像

    infiniflow/ragflow

    82,922在 GitHub 上查看↗

    This project is a comprehensive retrieval-augmented generation platform designed for building, managing, and deploying knowledge-based AI applications. It provides a unified environment for organizing datasets, configuring conversational chat assistants, and developing autonomous agents that execute multi-step reasoning workflows. By integrating document intelligence with advanced retrieval pipelines, the platform enables the creation of grounded, verifiable responses supported by traceable citations. The platform distinguishes itself through deep document understanding and sophisticated know

    This platform provides a complete RAG pipeline with built-in document parsing, multi-source ingestion, and vector-based retrieval, making it a comprehensive solution for indexing and querying internal documentation.

    PythonDocument Parsing Pipelines
    在 GitHub 上查看↗82,922
  • swirlai/swirl-searchswirlai 的头像

    swirlai/swirl-search

    3,032在 GitHub 上查看↗

    Swirl Search is an AI-powered middleware engine designed for federated enterprise search and retrieval-augmented generation. It functions as a unified platform that aggregates search results from multiple internal data sources and applications simultaneously, allowing users to query disparate knowledge bases without requiring data migration or indexing. The platform distinguishes itself by using a modular pipeline architecture to transform queries and synthesize information into grounded natural language answers. It employs semantic re-ranking and vector similarity to normalize results across

    This is a self-hostable enterprise search platform that provides a RAG pipeline and federated search across multiple data sources, directly addressing the need for querying internal documentation without moving data.

    PythonRetrieval-Augmented GenerationExternal Data Connectors
    在 GitHub 上查看↗3,032
  • cinnamon/kotaemonCinnamon 的头像

    Cinnamon/kotaemon

    25,139在 GitHub 上查看↗

    Kotaemon is an orchestration framework designed for building modular, agentic workflows that integrate document processing, retrieval-augmented generation, and multi-step reasoning. It provides a comprehensive platform for developing document-based question answering systems, allowing users to chain language models, prompt templates, and external tools into complex, automated pipelines. The system distinguishes itself through a highly modular architecture that emphasizes component-based composition and schema-driven data exchange. It supports autonomous agents capable of decomposing complex q

    Kotaemon is an orchestration framework specifically built for RAG-based document question answering, providing the necessary pipelines for document parsing, retrieval, and agentic workflows to index and query internal data.

    PythonDocument Parsing PipelinesPrivate Data HostingRetrieval-Augmented Generation Frameworks
    在 GitHub 上查看↗25,139
  • stangirard/quivrStanGirard 的头像

    StanGirard/quivr

    39,167在 GitHub 上查看↗

    Quivr is a framework for building retrieval-augmented generation pipelines that connect large language models to custom knowledge bases. It serves as a generative AI integration layer that abstracts the process of transforming diverse document sources into searchable context for AI responses. The project orchestrates the end-to-end flow between document ingestion, vector storage management, and model provider interfaces. It features a vector-store-agnostic retrieval system and a modular API layer that allows for flexible switching between different generative model providers. The system cove

    Quivr is a framework for building RAG pipelines that handles document ingestion and vector retrieval, providing the core infrastructure needed to index and query internal data for AI-powered search.

    PythonRetrieval Augmented Generation
    在 GitHub 上查看↗39,167
  • yichuan-w/leannyichuan-w 的头像

    yichuan-w/LEANN

    11,985在 GitHub 上查看↗

    LEANN is a framework for local retrieval augmented generation and vector indexing. It functions as a system for building local knowledge bases and source code search engines that combine large language models with retrieved private data to generate context-aware responses. The project distinguishes itself through a vision-model based document layout extractor for parsing complex PDF figures and diagrams, and a source code search engine that employs structure-aware chunking to preserve function and class boundaries. It also implements the Model Context Protocol to integrate real-time data sour

    LEANN is a framework for building local RAG and retrieval systems that supports document parsing and vector indexing, providing the core capabilities needed to index and query private data.

    PythonRetrieval Augmented Generation
    在 GitHub 上查看↗11,985
  • hkuds/lightragHKUDS 的头像

    HKUDS/LightRAG

    36,651在 GitHub 上查看↗

    LightRAG is a graph-based retrieval framework designed to build retrieval-augmented generation pipelines. It structures unstructured text into knowledge graphs, enabling multi-hop reasoning and complex query synthesis across large document collections. By integrating dense vector embeddings with structured knowledge graphs, the system facilitates both similarity-based and relationship-aware information retrieval. The framework distinguishes itself through a dual-level retrieval strategy that combines low-level keyword matching with high-level semantic graph traversal to capture both specific

    LightRAG is a specialized framework for building RAG pipelines that combine vector embeddings with knowledge graphs, providing the core retrieval and indexing capabilities needed for an enterprise search system.

    PythonRetrieval Augmented Generation Pipelines
    在 GitHub 上查看↗36,651
  • timescale/pgaitimescale 的头像

    timescale/pgai

    5,802在 GitHub 上查看↗

    pgai is a PostgreSQL AI toolkit and framework designed to integrate large language models and vector embeddings directly into a database. It serves as a bridge for executing machine learning model requests and performing text-to-SQL translations within standard database queries. The project provides an automated vector embedding pipeline that handles the loading, parsing, and chunking of text from tables and unstructured documents. This system utilizes a background worker to synchronize embeddings automatically as source data changes and includes specialized tools for building retrieval-augme

    This toolkit provides the essential RAG pipeline, vector embedding, and document parsing capabilities required to build an enterprise search system directly within PostgreSQL, though it functions as a framework for building such a system rather than a pre-packaged, standalone search application.

    PLpgSQLRetrieval-Augmented GenerationDocument and Unstructured Extraction
    在 GitHub 上查看↗5,802
  • labring/fastgptlabring 的头像

    labring/FastGPT

    27,132在 GitHub 上查看↗

    FastGPT is a comprehensive platform for building, deploying, and managing context-aware artificial intelligence applications. It provides a unified environment that integrates custom data sources with language models, utilizing a retrieval-augmented generation engine to ground responses in accurate, domain-specific information. The system is designed for enterprise-scale use, featuring multi-tenant architecture, administrative controls, and secure authentication protocols including OAuth 2.0 and custom single sign-on integration. The platform distinguishes itself through a visual, node-based

    FastGPT is a self-hostable RAG platform that provides the necessary document parsing, vector database integration, and access control features to build an enterprise-grade internal knowledge retrieval system.

    TypeScriptRetrieval-Augmented Generation Frameworks
    在 GitHub 上查看↗27,132
  • deepset-ai/haystackdeepset-ai 的头像

    deepset-ai/haystack

    24,253在 GitHub 上查看↗

    Haystack is an orchestration framework designed for building complex search and generative AI pipelines. It functions as an agentic workflow engine, enabling the construction of automated sequences that allow AI agents to perform multi-step reasoning and data analysis. The framework utilizes a modular, component-based architecture that connects processing steps into directed acyclic graphs. By employing a provider-agnostic integration layer, it decouples core logic from specific external AI services and vector databases, allowing for the flexible exchange of underlying technologies. This desi

    Haystack is a modular framework for building custom RAG pipelines and search systems, providing the necessary components for document indexing and retrieval even though it requires assembly rather than being a pre-packaged enterprise application.

    MDXVector Database Integrations
    在 GitHub 上查看↗24,253
  • nomic-ai/gpt4allnomic-ai 的头像

    nomic-ai/gpt4all

    77,375在 GitHub 上查看↗

    GPT4All is a cross-platform runtime environment designed to execute large language models directly on local consumer hardware. By leveraging an optimized C++ inference backend, it enables private, offline AI interactions without requiring an internet connection or external cloud services. The project provides a comprehensive ecosystem for managing the entire model lifecycle, including discovery, downloading, and configuration of local weights. What distinguishes the platform is its integrated retrieval-augmented generation engine, which allows users to index local documents into semantic vect

    This is a local-first AI application that includes a built-in RAG pipeline and document indexing capabilities, making it a suitable tool for querying internal data on a private, self-hosted basis.

    C++Retrieval Augmented Generation
    在 GitHub 上查看↗77,375
  • yeachan-heo/oh-my-codexYeachan-Heo 的头像

    Yeachan-Heo/oh-my-codex

    30,984在 GitHub 上查看↗

    oh-my-codex is an AI coding workflow orchestrator and a retrieval augmented generation documentation assistant. It manages complex programming tasks through a structured sequence of planning, execution, and verification phases, while providing tools for querying and translating technical documentation. The project utilizes Git worktrees to isolate parallel coding sessions, ensuring that concurrent tasks remain independent. It integrates a vector-store knowledge base to index documents into embeddings, enabling semantic search and factual context retrieval across multiple languages. The syste

    This is an AI-powered documentation assistant that features vector indexing and RAG capabilities for technical knowledge retrieval, though it is more specialized toward coding workflows than a general-purpose enterprise search platform.

    TypeScriptRetrieval-Augmented Generation
    在 GitHub 上查看↗30,984
  • typesense/typesensetypesense 的头像

    typesense/typesense

    25,254在 GitHub 上查看↗

    Typesense is a distributed search engine designed to provide sub-millisecond query latency across massive datasets. It functions as both a high-performance indexing and retrieval engine and a comprehensive search experience platform, offering built-in typo tolerance and tools for managing relevance through synonym configuration, result curation, and complex filtering. The platform distinguishes itself by utilizing in-memory indexing to maintain high-throughput data retrieval and integrating vector database capabilities to support semantic similarity searches. It ensures data consistency and h

    Typesense is a high-performance, self-hostable search engine that includes native vector database capabilities and semantic search, making it a strong foundation for building an enterprise retrieval system, though it requires additional orchestration to implement a full RAG pipeline and document parsing.

    C++Vector Databases
    在 GitHub 上查看↗25,254
  • khoj-ai/khojkhoj-ai 的头像

    khoj-ai/khoj

    35,163在 GitHub 上查看↗

    Khoj is a self-hosted artificial intelligence platform designed for personal knowledge management and semantic information retrieval. It functions as a private assistant that indexes your local documents, notes, and external workspaces, allowing you to interact with your data through natural language queries and conversational chat. By maintaining a local-first architecture, the system ensures that your information remains under your control while providing context-aware responses grounded in your personal knowledge base. The platform distinguishes itself through a modular, cross-platform int

    Khoj is a self-hosted AI platform that provides semantic search and RAG capabilities over personal and workspace documents, making it a functional tool for internal data retrieval despite its primary focus on individual knowledge management.

    PythonRetrieval-Augmented Generation
    在 GitHub 上查看↗35,163
  • othmanadi/planning-with-filesOthmanAdi 的头像

    OthmanAdi/planning-with-files

    14,139在 GitHub 上查看↗

    Planning with files is an enterprise knowledge graph platform designed to transform unstructured organizational data into a searchable, interconnected network. By utilizing a graph-based retrieval-augmented generation engine, the system grounds language model outputs in verified internal data, ensuring that responses are explainable, traceable, and free from hallucinations. The platform distinguishes itself through a focus on data sovereignty and secure, private infrastructure deployment. It enables organizations to maintain full control over sensitive information by processing data locally o

    This platform functions as an enterprise knowledge graph and retrieval-augmented generation engine designed to index and query internal data, aligning with the core requirements for an enterprise search and retrieval system.

    PythonDocument and Unstructured Extraction
    在 GitHub 上查看↗14,139
  • microsoft/graphragmicrosoft 的头像

    microsoft/graphrag

    33,792在 GitHub 上查看↗

    GraphRAG is a data processing pipeline and retrieval engine designed to transform unstructured text into interconnected knowledge graphs. By utilizing language models to extract entities and relationships, it builds structured representations of information that enable context-aware retrieval for downstream applications. The system distinguishes itself through hierarchical graph clustering and large-scale data synthesis, which organize massive document corpora into multi-level structures. This approach allows for both vector-based semantic searches and graph-based traversals, providing a comp

    This is a specialized RAG framework that builds knowledge graphs for advanced retrieval, providing the core indexing and query capabilities needed for enterprise documentation even though it functions as a pipeline engine rather than a turnkey, all-in-one search application.

    PythonGraph-Based Retrieval AugmentationGraph-Based Retrieval EnginesContext-Aware Retrieval
    在 GitHub 上查看↗33,792
  • meilisearch/meilisearchmeilisearch 的头像

    meilisearch/meilisearch

    58,118在 GitHub 上查看↗

    Meilisearch is a Rust-based search engine providing typo-tolerant full-text and vector-based semantic search with real-time conversational capabilities.

    Meilisearch is a high-performance search engine that supports vector-based semantic search and document indexing, making it a capable foundation for building an internal retrieval system, though it lacks built-in RAG pipelines and native role-based access control.

    RustDeveloper-Focused Search ToolsDocument Indexing EnginesFinite State Transducers
    在 GitHub 上查看↗58,118
  • thinkany-ai/rag-searchthinkany-ai 的头像

    thinkany-ai/rag-search

    1,179在 GitHub 上查看↗

    This project provides a search service designed to retrieve and rerank web content for use in large language model applications. It functions as a retrieval augmented search engine that processes natural language queries to fetch contextually relevant information from external web sources. The system distinguishes itself through a combination of semantic retrieval and precision-focused reranking. It converts user queries into high-dimensional embeddings to perform similarity searches across indexed collections, then refines these results by passing candidate pairs through a secondary model to

    This repository provides a RAG-based search API designed for indexing and querying documents, serving as a core component for building an enterprise search and retrieval system.

    PythonRetrieval Augmented Generation
    在 GitHub 上查看↗1,179
  • microsoft/pike-ragmicrosoft 的头像

    microsoft/PIKE-RAG

    2,388在 GitHub 上查看↗

    Pike-RAG is a framework designed for industrial-grade language model applications that require high factual accuracy and logical consistency. It functions as a platform for orchestrating multi-agent systems and implementing rationale-augmented generation, ensuring that model outputs are grounded in specialized domain knowledge rather than relying solely on internal training data. The system distinguishes itself through its ability to decompose complex, high-level queries into atomic tasks that are executed by specialized autonomous agents. By enforcing explicit logical reasoning steps before

    This project provides a specialized RAG pipeline and knowledge extraction framework designed for domain-specific retrieval, making it a functional tool for building an enterprise search system despite its focus on rationale-augmented generation.

    PythonRetrieval-Augmented Generation Frameworks
    在 GitHub 上查看↗2,388
  • 1517005260/graph-rag-agent1517005260 的头像

    1517005260/graph-rag-agent

    2,240在 GitHub 上查看↗

    This project is a comprehensive framework for constructing, managing, and evaluating knowledge graphs through multi-agent reasoning and deep search capabilities. It provides an end-to-end pipeline that ingests multi-format documents, extracts entities and relationships based on configurable schemas, and maintains structured knowledge bases to support evidence-based retrieval. The system distinguishes itself through its multi-agent orchestration, which decomposes complex queries into parallel research steps and synthesizes long-form reports. It leverages advanced graph-based techniques, includ

    This project provides a specialized RAG pipeline focused on knowledge graph construction and reasoning, serving as a functional tool for indexing and querying private data despite lacking built-in role-based access control or a broad suite of enterprise connectors.

    PythonBusiness Knowledge AgentsGraph RAG FrameworksGraph Retrieval Augmented Generation
    在 GitHub 上查看↗2,240
  • mixedbread-ai/mgrepmixedbread-ai 的头像

    mixedbread-ai/mgrep

    3,289在 GitHub 上查看↗

    mgrep is an LLM-powered semantic search engine and local file indexer designed to retrieve information from local directories and web content using natural language queries. It functions as a semantic document retriever that uses meaning and context rather than exact keyword matches to locate relevant data. The tool distinguishes itself by combining local file indexing with real-time web content retrieval to synthesize comprehensive answers. It employs retrieval-augmented generation to transform retrieved snippets from both local and remote sources into direct, concise responses. The system

    This is a semantic search and retrieval engine that supports local document indexing and RAG-based response synthesis, making it a capable tool for querying internal data despite lacking enterprise-grade features like role-based access control.

    TypeScriptSemantic Search EnginesFile and Web Search ToolsHybrid Retrieval
    在 GitHub 上查看↗3,289
  • opensemanticsearch/open-semantic-searchopensemanticsearch 的头像

    opensemanticsearch/open-semantic-search

    1,181在 GitHub 上查看↗

    Open Semantic Search is an open-source enterprise discovery platform designed to index, analyze, and explore large, diverse document collections. It functions as a comprehensive search engine and analytics suite that transforms unstructured data into structured information through automated processing pipelines. The platform distinguishes itself by integrating semantic exploration with traditional retrieval methods. It utilizes knowledge graph entity linking and thesaurus-driven query expansion to connect related concepts, allowing users to navigate datasets beyond simple keyword matching. Th

    Open Semantic Search is a comprehensive enterprise discovery and indexing platform that provides the necessary document parsing, multi-source ingestion, and search-based navigation required for internal data retrieval, though it focuses more on knowledge graph and semantic analysis than on a modern vector-based RAG pipeline.

    ShellEnterprise SearchSemantic Search EnginesEnterprise Discovery Platforms
    在 GitHub 上查看↗1,181
  • tobi/qmdtobi 的头像

    tobi/qmd

    9,498在 GitHub 上查看↗

    qmd is a local semantic search engine and RAG knowledge base indexer that functions as a Model Context Protocol server. It converts local documents, markdown files, and codebases into a searchable database to provide retrieval augmented generation capabilities for AI agents. The system exposes its search and retrieval tools via stdio or HTTP. It utilizes local model files for embeddings and reranking, supporting query expansion across multiple languages. The project employs abstract syntax tree based chunking to split source code at function and class boundaries. It implements hybrid vector-

    This is a local semantic search and RAG indexing tool that provides the core retrieval capabilities for internal documentation, though it lacks enterprise-grade features like role-based access control and multi-source connectors.

    TypeScriptLocal Knowledge Base IndexersSemantic Search EnginesAI-Workflow Code Index Builders
    在 GitHub 上查看↗9,498
  • knowledgecanvas/knowledgeKnowledgeCanvas 的头像

    KnowledgeCanvas/knowledge

    1,458在 GitHub 上查看↗

    Knowledge is a tool for saving, searching, accessing, exploring and chatting with all of your favorite websites, documents and files.

    This tool provides a platform for indexing and chatting with personal or organizational documents, offering the core retrieval and AI-interaction capabilities required for an internal knowledge system.

    TypeScriptKnowledge Management
    在 GitHub 上查看↗1,458
一览前 10 名对比
仓库Star 数语言许可证最后推送
onyx-dot-app/onyx17.5KPythonother2026年2月20日
truefoundry/cognita4.3KPythonapache-2.02025年11月26日
danswer-ai/danswer30.6KPythonNOASSERTION2026年6月26日
hkuds/rag-anything21.4KPythonMIT2026年6月15日
modsetter/surfsense14.8KPythonApache-2.02026年6月15日
quivrhq/quivr39.2KPythonNOASSERTION2025年7月9日
mintplex-labs/anything-llm61.7KJavaScriptMIT2026年6月16日
zylon-ai/private-gpt57.3KPythonApache-2.02026年6月16日
asyncfuncai/deepwiki-open14.4KPythonmit2026年1月25日
infiniflow/ragflow82.9KPythonApache-2.02026年6月16日

Related searches

  • 基于图结构的 RAG 框架
  • Knowledge base software
  • 用于 RAG 的混合检索引擎
  • 支持知识库
  • 文本实体提取工具包
  • 用于构建 RAG 流水线的框架
  • Multimodal retrieval system
  • 用于 RAG 检索的重排序库