awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
timescale avatar

timescale/pgaiArchived

0
View on GitHub↗
5,802 estrellas·311 forks·PLpgSQL·PostgreSQL·19 vistas

Pgai

pgai es un kit de herramientas y framework de IA para PostgreSQL diseñado para integrar modelos de lenguaje de gran tamaño (LLM) y embeddings vectoriales directamente en la base de datos. Actúa como un puente para ejecutar solicitudes de modelos de machine learning y realizar traducciones de texto a SQL dentro de consultas estándar de base de datos.

El proyecto proporciona un pipeline automatizado de embeddings vectoriales que gestiona la carga, el análisis y la fragmentación de texto desde tablas y documentos no estructurados. Este sistema utiliza un worker en segundo plano para sincronizar los embeddings automáticamente a medida que cambian los datos de origen e incluye herramientas especializadas para crear aplicaciones de generación aumentada por recuperación (RAG) y motores de búsqueda semántica.

El kit de herramientas cubre amplias áreas de capacidad, incluyendo el procesamiento de datos no estructurados con OCR, la creación de catálogos semánticos para mapear esquemas de bases de datos a lenguaje natural, y la implementación de búsquedas de similitud de alto rendimiento mediante indexación vectorial y reordenamiento de resultados. También permite el enriquecimiento de datos, la clasificación y la moderación de contenido llamando a modelos externos mediante SQL.

Features

  • AI Model Integrations - Integrates external AI models directly into PostgreSQL via SQL queries for natural language processing and embedding generation.
  • Database AI Toolkits - Provides a comprehensive set of tools to integrate LLMs and vector embeddings directly into PostgreSQL.
  • AI Model Integrations - Integrates external machine learning models directly into PostgreSQL queries for data enrichment and classification.
  • SQL-Based Machine Learning - Enables executing machine learning model requests and inference directly within standard SQL queries.
  • Standard RAG Development - Builds retrieval augmented generation pipelines that combine database retrieval with language models for grounded responses.
  • Retrieval-Augmented Generation - Implements the full retrieval-augmented generation pipeline by combining semantic search results with language model prompts.
  • SQL-Based Model Invocations - Allows executing model requests for text generation, classification, and moderation directly within SQL queries.
  • RAG Pipelines - Implements workflows that augment language model outputs by retrieving and integrating relevant external database data.
  • RAG Application Frameworks - Provides a framework to build RAG applications by combining retrieved database context with model prompts.
  • RAG Frameworks - Provides a framework for building retrieval-augmented generation applications using database context and LLM prompts.
  • SQL-Based Model Invocations - Executes requests to external machine learning models directly from within data queries.
  • Text-to-SQL Translators - Translates natural language user queries into executable SQL statements using semantic catalogs and schema descriptions.
  • Semantic Catalogs - Maintains semantic catalogs that map database objects to natural language descriptions to improve text-to-SQL accuracy.
  • Vector Embeddings - Generates numerical vector representations of database tables and files to enable semantic search.
  • Semantic Vector Search - Retrieves relevant data by calculating the mathematical distance between query and document embeddings.
  • Vector Similarity Search - Performs high-dimensional similarity searches to retrieve data based on semantic meaning rather than keywords.
  • Automatic Vector Embeddings - Automatically converts database content into vector representations using a background worker process.
  • Embedding Synchronization - Automatically synchronizes vector embeddings as source table data changes using state-based tracking.
  • Document Chunking and Embedding Pipelines - Provides automated pipelines that handle the full flow of chunking, embedding, and storing document data.
  • Embedding Pipelines - Implements modular pipelines that automate the loading, parsing, and formatting of data into vector embeddings.
  • Vector-Augmented Queries - Combines semantic vector search with model prompting within standard database queries to build knowledge-aware applications.
  • Semantic Search Engines - Implements a semantic search engine that retrieves information based on conceptual meaning using vector embeddings.
  • Continuous Sync Engines - Implements engines that automatically synchronize vector embeddings as the underlying source data changes.
  • In-Database Model Invocation - Enables executing external machine learning model requests and text-to-SQL translations directly within standard database queries.
  • Vector Indexing - Creates and manages indexes optimized for high-dimensional vector data to support semantic search.
  • Embedding Generation - Provides an automated pipeline to convert database content into high-dimensional vector representations using external workers.
  • Recursive Text Splitting - Implements recursive splitting functions to divide large bodies of text into chunks for AI model consumption.
  • Text Chunks - Splits long text into smaller segments using configurable algorithms and metadata injection for RAG context windows.
  • Result Reranking - Scores and reorders search results against a query to improve precision and relevance.
  • Embedding Status Monitors - Tracks the status of asynchronous embedding jobs, including pending items and processing states.
  • Data Processing - Uses large language models to enrich relational data through automated summarization, categorization, and content moderation.
  • Document Parsing and Extraction - Extracts text from unstructured PDFs and images using OCR to prepare content for vectorization and LLM ingestion.
  • Schema Description Generators - Generates natural language descriptions of database objects to help AI models understand technical schemas.
  • Embedding Synchronization Schedulers - Automates the periodic processing of updated data to synchronize embeddings via a background job system.
  • Document and Unstructured Extraction - Extracts text from PDFs and images using OCR and layout-aware parsing to prepare unstructured content for embedding.
  • Batch Embedding Management - Manages large-scale batch processing of embeddings with built-in resilience against failures and API rate limits.
  • Query Performance Tuning - Optimizes vector query performance by creating specialized indexes on embedding columns to reduce latency.
  • Semantic Knowledge Base Search - Retrieves relevant context and domain knowledge from the database using natural language queries and embeddings.
  • Background Job Processing - Provides a background worker system to process large-scale embedding tasks and synchronization jobs outside the main request flow.
  • Catalog Management - Enables the management of multiple independent semantic catalogs and embedding configurations within one environment.
  • Spatial and Vector Data - Simplifies creating and synchronizing vector embeddings.

Historial de estrellas

Gráfico del historial de estrellas de timescale/pgaiGráfico del historial de estrellas de timescale/pgai

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Preguntas frecuentes

¿Qué hace timescale/pgai?

pgai es un kit de herramientas y framework de IA para PostgreSQL diseñado para integrar modelos de lenguaje de gran tamaño (LLM) y embeddings vectoriales directamente en la base de datos. Actúa como un puente para ejecutar solicitudes de modelos de machine learning y realizar traducciones de texto a SQL dentro de consultas estándar de base de datos.

¿Cuáles son las características principales de timescale/pgai?

Las características principales de timescale/pgai son: AI Model Integrations, Database AI Toolkits, SQL-Based Machine Learning, Standard RAG Development, Retrieval-Augmented Generation, SQL-Based Model Invocations, RAG Pipelines, RAG Application Frameworks.

¿Qué alternativas de código abierto existen para timescale/pgai?

Las alternativas de código abierto para timescale/pgai incluyen: casibase/casibase — Casibase is an open-source platform that orchestrates multi-turn conversations with large language models and manages… chonkie-inc/chonkie — Chonkie is a text chunking library designed for retrieval-augmented generation pipelines. It functions as a semantic… docker/genai-stack — This project is a containerized development stack and application framework for building retrieval-augmented… datawhalechina/all-in-rag — This project is a retrieval augmented generation framework designed to build pipelines that connect unstructured data… ravendb/ravendb — RavenDB is a multi-model NoSQL document database designed for high-performance, ACID-compliant data storage. It… azure-samples/azure-search-openai-demo — This project is a reference implementation and application template for Retrieval-Augmented Generation (RAG). It…

Alternativas open-source a Pgai

Proyectos open-source similares, clasificados según cuántas características comparten con Pgai.
  • casibase/casibaseAvatar de casibase

    casibase/casibase

    4,443Ver en GitHub↗

    Casibase is an open-source platform that orchestrates multi-turn conversations with large language models and manages retrieval-augmented knowledge bases from a single interface. It provides a unified system for connecting to over 30 AI model providers, ingesting documents into vector embeddings for semantic search, and running autonomous agent loops that can drive a browser, search the web, execute commands, and integrate with external tools. The platform distinguishes itself by combining AI conversation management with infrastructure and application orchestration capabilities. It includes a

    Goa2aagentagi
    Ver en GitHub↗4,443
  • chonkie-inc/chonkieAvatar de chonkie-inc

    chonkie-inc/chonkie

    4,170Ver en GitHub↗

    Chonkie is a text chunking library designed for retrieval-augmented generation pipelines. It functions as a semantic text splitter and RAG ingestion pipeline, transforming raw text into embedded segments for storage in vector databases. The project distinguishes itself through specialized splitting strategies, including an AST-based code splitter for preserving logical boundaries in source code and a semantic text splitter that uses embedding models to determine boundaries based on meaning. It also provides a vector database ingestor to automate the generation of embeddings and their export t

    Pythonaichonkiechunker
    Ver en GitHub↗4,170
  • docker/genai-stackAvatar de docker

    docker/genai-stack

    5,333Ver en GitHub↗

    This project is a containerized development stack and application framework for building retrieval-augmented generation systems. It provides a dockerized AI sandbox that integrates local model runtimes, knowledge graphs, and vector stores to enable the creation of contextual chatbots. The stack is distinguished by its graph-based vector store, which combines structured knowledge graphs with vector indices for both semantic and structural data retrieval. It allows for local model hosting with CPU or GPU acceleration, enabling generative tasks without reliance on external cloud APIs. The frame

    Python
    Ver en GitHub↗5,333
  • datawhalechina/all-in-ragAvatar de datawhalechina

    datawhalechina/all-in-rag

    3,989Ver en GitHub↗

    This project is a retrieval augmented generation framework designed to build pipelines that connect unstructured data and knowledge graphs with large language models. It functions as a vector database orchestrator for indexing text and multimodal content, as well as a system for translating natural language queries into structured database commands. The framework integrates a hybrid retrieval engine that combines dense vector search with sparse keyword matching to increase the precision of retrieved contexts. It further enhances reasoning and relationship mapping through a graph-augmented ret

    Pythonaideepseekembedding
    Ver en GitHub↗3,989
Ver las 30 alternativas a Pgai→