awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

11 repositorios

Awesome GitHub RepositoriesKnowledge Base Construction

End-to-end processes for parsing documents, generating embeddings, and storing chunks for semantic retrieval.

Distinct from Index Construction: Covers the full pipeline from parsing to storage, whereas index construction focuses only on populating the structure.

Explore 11 awesome GitHub repositories matching data & databases · Knowledge Base Construction. Refine with filters or upvote what's useful.

Awesome Knowledge Base Construction GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • netease-youdao/qanythingAvatar de netease-youdao

    netease-youdao/QAnything

    14,020Ver en GitHub↗

    QAnything is a retrieval-augmented generation application framework and self-hosted AI interface. It functions as a system that combines a vector database knowledge base, a document parsing service, and a hybrid search engine to generate answers based on private user data. The project features a modular pipeline architecture that allows users to independently replace components such as parsers, embedding models, and reranking engines. It supports local-first model deployment and offline operation to ensure data privacy, and includes a two-stage retrieval pipeline that merges dense vector embe

    Implements end-to-end processes for parsing documents, generating embeddings, and storing chunks for semantic retrieval.

    Python
    Ver en GitHub↗14,020
  • tporadowski/redisAvatar de tporadowski

    tporadowski/redis

    9,987Ver en GitHub↗

    Redis is a high-performance in-memory key-value store that functions as a distributed cache, message broker, and NoSQL database. It provides sub-millisecond read and write access to data stored in RAM and can operate as a vector database for indexing high-dimensional embeddings. The system supports a wide range of data storage and synchronization primitives, including the management of strings, hashes, lists, sets, and JSON documents. It enables real-time data operations through atomic transactions, hybrid persistence using snapshots and append-only logs, and high-availability configurations

    Parses documents and creates vector embeddings to build searchable knowledge bases for semantic search.

    Credisredis-for-windowsredis-msi-installer
    Ver en GitHub↗9,987
  • yusufkaraaslan/skill_seekersAvatar de yusufkaraaslan

    yusufkaraaslan/Skill_Seekers

    9,641Ver en GitHub↗

    Skill Seekers is a toolset for generating large language model knowledge bases, featuring a multi-source content scraper and a dedicated RAG data pipeline. It extracts technical data from documentation, code, and video to create structured assets and configuration files for AI-powered IDE extensions. The project distinguishes itself through the ability to transform raw data into polished tutorials and specialized skills for AI plugin marketplaces. It utilizes abstract syntax tree parsing and optical character recognition to analyze GitHub repositories, PDFs, and video frames, converting these

    Converts documentation and diverse data sources into structured formats for retrieval pipelines and vector databases.

    Pythonai-toolsast-parserautomation
    Ver en GitHub↗9,641
  • liaokongvfx/langchain-chinese-getting-started-guideAvatar de liaokongVFX

    liaokongVFX/LangChain-Chinese-Getting-Started-Guide

    9,039Ver en GitHub↗

    This project is a collection of tutorials and guides for building large language model applications using the LangChain framework, written in Chinese. It serves as a learning resource for developing software that integrates language models with memory and chain-based logic. The resource provides specific walkthroughs for implementing retrieval augmented generation systems using vector stores and document loaders. It includes guides on creating autonomous agents that dynamically select and execute external tools, as well as tutorials for translating plain text queries into executable database

    Walks through the full pipeline of parsing documents and generating embeddings for semantic retrieval.

    Ver en GitHub↗9,039
  • 53ai/53aihubAvatar de 53AI

    53AI/53AIHub

    9,025Ver en GitHub↗

    53AIHub is a centralized orchestration platform for deploying and managing AI agents and prompts across multiple large language model providers. It functions as a multi-model AI gateway and an operation portal for AI services, providing a unified interface to coordinate agents and prompts from various external platforms. The project distinguishes itself as a white-label AI portal designed for self-hosted infrastructure, allowing for full control over operational data on private servers or containers. It includes a comprehensive AI SaaS administration layer with a multi-tenant subscription eng

    Implements end-to-end processes for parsing documents and generating embeddings for semantic retrieval.

    Gocozedifyfastgpt
    Ver en GitHub↗9,025
  • intel/ipex-llmAvatar de intel

    intel/ipex-llm

    8,836Ver en GitHub↗

    Intel XPU LLM Acceleration Library is a toolkit designed to accelerate large language model inference and finetuning on Intel CPUs, GPUs, and NPUs. It provides a distributed inference engine for scaling models across multiple accelerators, a multimodal model runtime for vision and speech tasks, and a low-bit model quantization tool for converting weights into INT4, FP8, and GGUF formats. The project features a parameter-efficient finetuning framework that enables model adaptation using QLoRA and DPO on Intel hardware. It distinguishes itself by providing specialized optimizations for Intel XP

    Processes knowledge files to create searchable vector-based knowledge bases for question-answering tasks.

    Python
    Ver en GitHub↗8,836
  • wonderwhy-er/desktopcommandermcpAvatar de wonderwhy-er

    wonderwhy-er/DesktopCommanderMCP

    5,493Ver en GitHub↗

    DesktopCommanderMCP is a Model Context Protocol (MCP) server that gives AI agents direct access to local files, shell commands, and system processes through natural language instructions. It acts as a unified bridge between conversational commands and desktop operations, enabling an AI to translate plain English into file management, code editing, system command execution, data analysis, and software scaffolding tasks without needing its own API. The server exposes these capabilities as structured tools via the MCP protocol, so any compatible agent can interact with the local environment in a

    Creates, updates, and organizes local markdown notes and converts scattered data into structured living documents.

    TypeScriptagentaicode-analysis
    Ver en GitHub↗5,493
  • modelengine-group/nexentAvatar de ModelEngine-Group

    ModelEngine-Group/nexent

    5,265Ver en GitHub↗

    Nexent es un plano de control de IA empresarial y plataforma de orquestación de agentes LLM. Proporciona un entorno sin código (zero-code) para diseñar, desplegar y gestionar agentes de IA de producción a través de un framework de colaboración multi-agente que coordina agentes autónomos especializados utilizando protocolos de mensajería estandarizados. La plataforma integra el Model Context Protocol para conectar agentes con herramientas, plugins y servicios externos mediante una interfaz de comunicación universal. Destaca además con un gestor de base de conocimientos RAG dedicado que importa documentos no estructurados y utiliza búsqueda híbrida para proporcionar contexto fundamentado para las respuestas del modelo. El sistema cubre una amplia gama de capacidades, incluyendo control de acceso basado en roles multi-inquilino, interacción multimodal a través de texto, voz e imágenes, y recuperación vectorial híbrida. También incluye un mercado para la distribución y descubrimiento de agentes, junto con herramientas de observabilidad para capturar trazas de ejecución. La plataforma soporta despliegue seguro mediante empaquetado offline contenedorizado para infraestructura aislada (air-gapped).

    Parses and vectorizes various document formats into searchable knowledge bases with integrated access controls.

    Pythonagentagentic-aiagentic-framework
    Ver en GitHub↗5,265
  • tencentmusic/cube-studioAvatar de tencentmusic

    tencentmusic/cube-studio

    5,062Ver en GitHub↗

    Cube Studio es una plataforma MLOps nativa de la nube y un orquestador de IA basado en Kubernetes, diseñado para todo el ciclo de vida del machine learning. Proporciona un framework de entrenamiento distribuido para el ajuste fino de modelos a gran escala, un gestor de recursos GPU para virtualización de hardware y un orquestador de pipelines de ML que utiliza grafos acíclicos dirigidos visuales para gestionar flujos de trabajo de extremo a extremo. La plataforma se distingue por su servidor de inferencia LLM especializado, que soporta generación aumentada por recuperación (RAG) y la construcción de bases de conocimiento privadas. Cuenta con un sistema dedicado para el ajuste fino supervisado y aprendizaje por refuerzo de modelos de lenguaje grandes, complementado con herramientas visuales de búsqueda de hiperparámetros. El sistema cubre una amplia gama de capacidades operativas, incluyendo etiquetado de datos multimodales, pipelines de datos distribuidos y programación de cargas de trabajo en múltiples clústeres. También ofrece entornos de desarrollo interactivos basados en navegador, gestión de imágenes de contenedor y un registro de modelos para versionar y desplegar APIs de inferencia escalables con división de tráfico. La infraestructura incluye monitorización de salud del clúster integrada y control de acceso basado en roles con integración de inicio de sesión único (SSO).

    Integrates domain-specific data using embeddings and semantic retrieval to build private knowledge bases.

    Pythonaiaihubargo
    Ver en GitHub↗5,062
  • siteserver/cmsAvatar de siteserver

    siteserver/cms

    3,905Ver en GitHub↗

    This project is a .NET Core content management system and multi-site management platform designed to organize and publish structured digital content across independent websites from a centralized interface. It functions as both a headless CMS and a static site generator, rendering dynamic templates into HTML files to increase loading speed and scalability. The system integrates retrieval-augmented generation to transform website documents and content into searchable AI knowledge bases. It includes a visual AI workflow orchestrator to define the logic between user queries and large language mo

    Implements an end-to-end pipeline for parsing documents and generating embeddings to create searchable AI knowledge bases.

    JavaScriptc-sharpcmscontent-management-system
    Ver en GitHub↗3,905
  • badboysm890/claraverseAvatar de badboysm890

    badboysm890/ClaraVerse

    3,833Ver en GitHub↗

    ClaraVerse is a self-hosted orchestration platform for deploying and managing local language models, autonomous agents, and automated workflows on private infrastructure. It functions as a containerized backend manager that orchestrates services, databases, and model providers within local containers to maintain data sovereignty. The platform features a visual workflow builder with a drag-and-drop interface for designing complex parallel task sequences. It utilizes a multi-model abstraction layer to normalize interactions across diverse local and remote AI endpoints and includes a retrieval a

    Implements an end-to-end process for parsing documents and generating embeddings for semantic retrieval.

    Go
    Ver en GitHub↗3,833
  1. Home
  2. Data & Databases
  3. Index Construction
  4. Knowledge Base Construction

Explorar subetiquetas

  • Living Document ConvertersCreates, updates, and organizes local markdown notes and converts scattered data into structured living documents. **Distinct from Knowledge Base Construction:** Distinct from Knowledge Base Construction: focuses on converting scattered data into living documents rather than the full parsing-to-storage pipeline.
  • Multi-Tenant ImplementationsKnowledge base construction processes designed for multi-tenant environments with administrative access and data isolation. **Distinct from Knowledge Base Construction:** Focuses on the administrative and isolation requirements of multi-tenancy rather than just the technical parsing and embedding pipeline.