awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

19 repositorios

Awesome GitHub RepositoriesPostgreSQL Vector Stores

Configurations for using PostgreSQL with vector extensions as a knowledge base.

Explore 19 awesome GitHub repositories matching data & databases · PostgreSQL Vector Stores. Refine with filters or upvote what's useful.

Awesome PostgreSQL Vector Stores GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • zylon-ai/private-gptAvatar de zylon-ai

    zylon-ai/private-gpt

    57,278Ver en GitHub↗

    This project is a privacy-first backend service designed to facilitate retrieval-augmented generation by processing local documents into searchable vector representations. It provides a modular architecture that allows users to ingest diverse file formats, manage document metadata, and perform semantic searches to provide context-aware responses for chat and completion requests. The system distinguishes itself through a database-agnostic abstraction layer that supports various storage backends, ranging from local disk storage to enterprise-grade vector databases. It offers flexible deployment

    Utilizes PostgreSQL as a scalable vector knowledge base through specialized configuration and dependency management.

    Python
    Ver en GitHub↗57,278
  • vectordotdev/vectorAvatar de vectordotdev

    vectordotdev/vector

    22,071Ver en GitHub↗

    Vector is a high-performance observability data pipeline designed to collect, transform, and route logs, metrics, and traces across distributed infrastructure. It functions as a modular engine that decouples data ingestion from processing and transmission, utilizing a component-based architecture to connect diverse sources to multiple destinations. The project distinguishes itself through a focus on reliability and flow control. It implements backpressure-aware data movement to prevent data loss during traffic spikes and utilizes disk-backed event buffering to ensure durability during network

    Writes logs, metrics, and traces into PostgreSQL databases using configurable batching and delivery guarantees.

    Rusteventsforwarderhacktoberfest
    Ver en GitHub↗22,071
  • openai/chatgpt-retrieval-pluginAvatar de openai

    openai/chatgpt-retrieval-plugin

    21,192Ver en GitHub↗

    This project is a retrieval-augmented generation pipeline designed for building custom ChatGPT plugins that allow language models to query private or professional documents. It implements a full retrieval workflow, from processing and indexing document chunks to retrieving relevant context for natural language queries. The system distinguishes itself through a hybrid retrieval approach that combines dense vector embeddings with sparse keyword matching, further refined by a two-stage semantic re-ranking process. It includes specialized data privacy tools for screening personally identifiable i

    Utilizes PostgreSQL with the pgvector extension to persist and manage document embeddings for retrieval.

    Pythonchatgptchatgpt-plugins
    Ver en GitHub↗21,192
  • activeloopai/hubAvatar de activeloopai

    activeloopai/Hub

    9,177Ver en GitHub↗

    Hub is a multimodal AI data lake and vector database designed for storing and querying embeddings, text, audio, and images. It functions as a dataset version control system and a machine learning data streaming engine to support large-scale model training. The system utilizes a serverless PostgreSQL vector store to index high-dimensional embeddings for semantic search. It provides a visual interface for inspecting multimodal datasets and viewing annotations such as bounding boxes and masks. The platform handles cloud-agnostic storage synchronization and implements lazy, compressed data strea

    Utilizes serverless PostgreSQL vector stores to index and store high-dimensional embeddings for semantic retrieval.

    C++
    Ver en GitHub↗9,177
  • arize-ai/phoenixAvatar de Arize-ai

    Arize-ai/phoenix

    8,605Ver en GitHub↗

    Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and monitor large language model applications. It serves as a prompt management system for versioning and testing templates, and as a self-hosted AI operations infrastructure for managing telemetry and experiments. The platform differentiates itself through a specialized embedding visualization tool used to detect data drift and optimize vector search. It provides a comprehensive evaluation suite that utilizes judge-based evaluators and ground-truth datasets to score model outputs, and

    Implements PostgreSQL data sinks to store telemetry and observability data with scalable ingestion.

    Jupyter Notebookagentsai-monitoringai-observability
    Ver en GitHub↗8,605
  • dimitri/pgloaderAvatar de dimitri

    dimitri/pgloader

    6,295Ver en GitHub↗

    pgloader is a command-line tool that automates the migration of data and schema from various source databases and file formats into PostgreSQL. It combines schema discovery, parallel data pipelines, and type casting into a single, declarative workflow, using PostgreSQL's COPY protocol for high-throughput bulk loading. The tool distinguishes itself by compiling a dedicated command language into concurrent reader-writer pipelines that handle schema introspection, data transformation, and error-resilient batch processing. It supports migrating entire databases from MySQL, MS SQL, SQLite, and Pos

    Automates migration of SQLite databases into PostgreSQL with schema discovery and index creation.

    Common Lispclozure-clcommon-lispcsv
    Ver en GitHub↗6,295
  • greptimeteam/greptimedbAvatar de GreptimeTeam

    GreptimeTeam/greptimedb

    5,968Ver en GitHub↗

    GreptimeDB is a distributed, open-source time-series database built for unified observability. It stores and queries metrics, logs, and traces together in a single columnar engine, supporting both SQL and PromQL for analysis. The database is designed as a Kubernetes-native operator with a decoupled compute and storage architecture, enabling horizontal scaling and multi-region deployment. What distinguishes GreptimeDB is its role as a multi-protocol ingestion gateway, accepting data through OpenTelemetry, Prometheus Remote Write, InfluxDB, Loki, Elasticsearch, Kafka, and MQTT protocols without

    Stores cluster metadata in PostgreSQL for production deployments.

    Rustanalyticscloud-nativedatabase
    Ver en GitHub↗5,968
  • dromara/lamp-cloudAvatar de dromara

    dromara/lamp-cloud

    5,752Ver en GitHub↗

    Lamp Cloud is a multi-tenant SaaS backend framework built on Java and Spring Cloud that provides a complete foundation for building enterprise-grade administration systems. Its core identity centers on supporting multiple tenant isolation strategies—including database-per-tenant, schema-per-tenant, and shared-table modes—that can be switched without altering business code, alongside a role-based access control system enforced at the gateway layer across all microservices. The framework distinguishes itself through comprehensive tenant lifecycle management tools that allow creating, configurin

    Uploads and retrieves files from FastDFS, MinIO, or other storage systems through a unified interface.

    Javaadmincloudeureka
    Ver en GitHub↗5,752
  • weblateorg/weblateAvatar de WeblateOrg

    WeblateOrg/weblate

    5,748Ver en GitHub↗

    Weblate is an open-source web-based translation management system that provides a collaborative platform for teams to review, suggest, and approve translations in real time. It functions as a continuous localization platform, automatically synchronizing translations with source code changes in version control repositories, and can be deployed either as a self-hosted server or through a managed cloud hosting service. The system integrates directly with Git hosting platforms like GitHub, GitLab, and Bitbucket, storing all translations in version control with individual translator attribution re

    Uses Django ORM with PostgreSQL and trigram extensions for translation data storage.

    Pythoncontinuous-localizationcrowdsourcingdcode-2025
    Ver en GitHub↗5,748
  • tortoise/tortoise-ormAvatar de tortoise

    tortoise/tortoise-orm

    5,582Ver en GitHub↗

    Tortoise ORM is an asynchronous object-relational mapper for Python that mirrors Django's model and queryset API while running on asyncio. It defines database tables as Python classes with typed fields and supports foreign key, many-to-many, and one-to-one relations, providing a chainable query API for filtering, annotating, grouping, and prefetching related objects without blocking the event loop. The ORM includes a built-in migration engine that detects model changes, generates migration files, and applies or reverts schema changes through a command-line tool. It connects to PostgreSQL, MyS

    Provides a Django-style ORM designed for async frameworks, enabling familiar data access in ASGI applications.

    Pythonasyncasynciomysql
    Ver en GitHub↗5,582
  • wger-project/wgerAvatar de wger-project

    wger-project/wger

    5,636Ver en GitHub↗

    wger is an open-source web application for fitness tracking, workout planning, and nutrition management. It provides a self-hosted platform where users can design weekly workout routines from a built-in exercise library, log their training progress, and plan daily meals using a food database with automatic nutritional calculations. The application supports multi-user accounts with credential-based login, passkey authentication, and third-party sign-in through OAuth providers. The platform includes a documented REST API that enables programmatic access to workout logs, meal plans, and user dat

    Uses Django's ORM to map Python objects to relational database tables with migration support.

    Pythondjangofitnessgym
    Ver en GitHub↗5,636
  • ckan/ckanAvatar de ckan

    ckan/ckan

    4,961Ver en GitHub↗

    CKAN is an open-source data management platform that provides the foundation for building data portals. It supports the full lifecycle of datasets—from creation and organization to publishing, cataloging with faceted search, and interactive data visualization—all through a web interface. The platform is built on a modular architecture that includes a plugin-based extensibility system, a harvesting framework for importing metadata from external sources, and a standardized RESTful JSON API for programmatic access to datasets and metadata. The web interface is rendered using the Jinja2 templatin

    Manages uploaded files in configurable storage (local filesystem or S3) while storing metadata in PostgreSQL.

    Pythonapicatalogckan
    Ver en GitHub↗4,961
  • jazzband/django-silkAvatar de jazzband

    jazzband/django-silk

    4,926Ver en GitHub↗

    Django Silk is a profiling and inspection toolset for Django applications designed to capture SQL queries, HTTP request data, and execution timing for diagnostics. It functions as a performance profiler and debugging middleware that records runtime execution data to provide a comprehensive overview of application behavior. The system includes a database profiler for identifying slow operations through detailed timing data and an HTTP request inspector for reviewing headers, bodies, and network traffic via a web interface. It allows for the reproduction of specific server requests through gene

    Stores all captured profiling records as Django model instances in a relational database.

    Python
    Ver en GitHub↗4,926
  • gomods/athensAvatar de gomods

    gomods/athens

    4,773Ver en GitHub↗

    Athens es un servidor proxy de módulos de Go y caché de dependencias que proporciona un sistema de almacenamiento persistente para dependencias de Go. Actúa como un espejo y almacén de datos para garantizar entornos de construcción reproducibles almacenando copias inmutables de paquetes externos, protegiendo contra eliminaciones o interrupciones en el origen. El proyecto destaca por servir como una puerta de enlace segura para el alojamiento de módulos de Go privados, utilizando tokens de autenticación, claves SSH y GitHub Apps para recuperar dependencias de sistemas de control de versiones privados. Además, permite el cumplimiento de las dependencias de software mediante el filtrado de solicitudes y el proxy de sumas de comprobación, lo que evita que los metadatos de los módulos privados se filtren a servidores públicos. El servidor admite una amplia gama de backends de almacenamiento, incluyendo disco local, bases de datos NoSQL y almacenes de objetos en la nube compatibles con S3. Incluye capacidades para el almacenamiento en caché de dependencias distribuido con bloqueo compartido para evitar descargas redundantes a través de múltiples instancias y proporciona herramientas para el prellenado de almacenamiento en entornos aislados (air-gapped). El servidor puede desplegarse a través de contenedores Docker, gráficos Helm de Kubernetes o varias plataformas en la nube gestionadas.

    Supports multiple configurable backends for storing downloaded module files, including S3 and local filesystems.

    Goathensdependenciesdependency-manager
    Ver en GitHub↗4,773
  • ruvnet/ruvectorAvatar de ruvnet

    ruvnet/ruvector

    4,253Ver en GitHub↗

    ruvector es un almacén de vectores y base de datos de grafos basado en Rust, diseñado para inferencia local y búsquedas de vecinos más cercanos. Utiliza una arquitectura de base de datos de grafos vectoriales y un índice de red neuronal de grafos para refinar los rankings de búsqueda mediante atención estructural. El sistema incluye un simulador de circuitos cuánticos acelerado por hardware para ejecutar simulaciones de vectores de estado y patrones de búsqueda complejos, junto con un motor de inferencia WebAssembly para ejecutar búsquedas vectoriales y modelos directamente en navegadores web. El proyecto emplea un formato de contenedor cognitivo que agrupa modelos, datos y un microkernel arrancable en un único binario para su despliegue. Incluye herramientas especializadas de configuración de modelos, como un método de consolidación de pesos para prevenir el olvido catastrófico y un mecanismo de adaptadores ligeros para la adaptación instantánea de pesos. El sistema cubre una amplia superficie de capacidades, incluyendo búsqueda vectorial acelerada por hardware, consultas de relaciones en grafos y análisis de documentos científicos para la extracción de LaTeX y MathML. También proporciona encadenamiento de testigos criptográficos para verificar mutaciones de datos, sincronización de metadatos basada en Raft para alta disponibilidad y compresión de datos de resolución escalonada para gestionar los costes de almacenamiento.

    Expands PostgreSQL capabilities with specialized SQL functions and self-learning vector search tools.

    Rust
    Ver en GitHub↗4,253
  • chonkie-inc/chonkieAvatar de chonkie-inc

    chonkie-inc/chonkie

    4,170Ver en GitHub↗

    Chonkie es una librería de fragmentación de texto (chunking) diseñada para pipelines de generación aumentada por recuperación (RAG). Funciona como un divisor de texto semántico y pipeline de ingesta RAG, transformando texto sin procesar en segmentos incrustados para su almacenamiento en bases de datos vectoriales. El proyecto se distingue por estrategias de división especializadas, incluyendo un divisor de código basado en AST para preservar límites lógicos en el código fuente y un divisor de texto semántico que utiliza modelos de embedding para determinar límites basados en el significado. También proporciona un ingestor de bases de datos vectoriales para automatizar la generación de embeddings y su exportación a varios almacenes. La librería cubre una amplia gama de capacidades, incluyendo el análisis de documentos mediante OCR y extracción de markdown, una variedad de métodos de división como conteo de tokens y segmentación jerárquica, y orquestación de flujos de trabajo a través de pipelines reutilizables. Admite una amplia gama de integraciones de almacenes vectoriales, incluyendo Qdrant, Milvus, Weaviate y Elasticsearch, así como la exportación de datos a JSON y datasets de Hugging Face. Los usuarios pueden ejecutar estas operaciones a través de una interfaz de línea de comandos o desplegar el sistema como un servicio API contenerizado.

    Saves processed text segments and vector embeddings into PostgreSQL using the pgvector extension.

    Pythonaichonkiechunker
    Ver en GitHub↗4,170
  • morpheus65535/bazarrAvatar de morpheus65535

    morpheus65535/bazarr

    4,070Ver en GitHub↗

    Bazarr is an automated subtitle management system and downloader designed to discover, acquire, and synchronize subtitles for movies and TV shows. It functions as a media library companion that integrates with external media managers and servers via APIs to track missing subtitles and ensure libraries are up to date. The project distinguishes itself through advanced media processing, using neural-network audio transcription to generate subtitles from audio tracks or translate foreign dialogue into English. It also features audio-based synchronization to align subtitle timing with video conten

    Utilizes a PostgreSQL database to store subtitle history and library state for improved scalability.

    Pythondownload-subtitlesepisodesmovie
    Ver en GitHub↗4,070
  • langroid/langroidAvatar de langroid

    langroid/langroid

    3,894Ver en GitHub↗

    Langroid is a multi-agent orchestration framework and tool integration suite designed for building complex AI applications. It serves as a multi-modal integration layer that connects diverse local and remote language models with an agentic retrieval-augmented generation system. The project distinguishes itself through a collaborative message-exchange paradigm, allowing specialized agents to delegate tasks hierarchically and coordinate via structured communication. It features an advanced state management system for conversational AI, including the ability to rewind and prune conversation hist

    Indexes document content in a PostgreSQL vector store for similarity searches.

    Pythonagentsaichatgpt
    Ver en GitHub↗3,894
  • apache/gravitinoAvatar de apache

    apache/gravitino

    2,866Ver en GitHub↗

    Gravitino is a federated metadata lake and unified data catalog designed to manage tables, files, and AI models across diverse data sources and cloud storage. It serves as a centralized interface for governing schemas, access controls, and tagging across relational databases, messaging queues, and object stores. The project distinguishes itself by unifying the management of AI assets, such as machine learning models and their version lineages, alongside traditional tabular data. It also implements the Iceberg REST specification to provide a standardized metadata server and proxy for lakehouse

    Governs schemas and tables within PostgreSQL databases, including the management of comments.

    Javaai-catalogdata-catalogdatalake
    Ver en GitHub↗2,866
  1. Home
  2. Data & Databases
  3. Database Management Systems
  4. Database Engines
  5. Vector Databases
  6. PostgreSQL Vector Stores

Explorar subetiquetas

  • PostgreSQL Data Sinks1 sub-etiquetaComponents for storing observability data in PostgreSQL databases with batching and delivery guarantees. **Distinct from PostgreSQL Vector Stores:** Distinct from PostgreSQL Vector Stores: focuses on general observability data storage rather than vector-specific extensions.
  • PostgreSQL Metadata Stores3 sub-etiquetasStoring cluster metadata in PostgreSQL for production deployments that integrate with existing database infrastructure. **Distinct from PostgreSQL Vector Stores:** Distinct from PostgreSQL Vector Stores: focuses on using PostgreSQL as a metadata store backend, not vector storage.
  • PostgreSQL Schema GovernanceGoverning schemas and tables within PostgreSQL databases. **Distinct from PostgreSQL Metadata Stores:** Focuses on the governance and administrative management of PostgreSQL schemas rather than using PostgreSQL as a metadata storage backend.