awesome-repositories.com
المدونة
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
timescale avatar

timescale/pgaiArchived

0
View on GitHub↗
5,802 نجوم·311 تفرعات·PLpgSQL·PostgreSQL·8 مشاهدات

Pgai

pgai هو مجموعة أدوات وإطار عمل لـ PostgreSQL مصمم لدمج نماذج اللغات الكبيرة وتضمينات المتجهات (vector embeddings) مباشرة داخل قاعدة البيانات. يعمل كجسر لتنفيذ طلبات نماذج تعلم الآلة وإجراء ترجمات النص إلى SQL ضمن استعلامات قاعدة البيانات القياسية.

يوفر المشروع خط أنابيب آلي لتضمين المتجهات يتولى تحميل وتحليل وتقسيم النصوص من الجداول والمستندات غير المهيكلة. يستخدم هذا النظام عاملاً في الخلفية لمزامنة التضمينات تلقائياً مع تغير البيانات المصدرية، ويتضمن أدوات متخصصة لبناء تطبيقات التوليد المعزز بالاسترجاع (RAG) ومحركات البحث الدلالي.

تغطي مجموعة الأدوات مجالات واسعة تشمل معالجة البيانات غير المهيكلة باستخدام OCR، وإنشاء فهارس دلالية لربط مخططات قاعدة البيانات باللغة الطبيعية، وتنفيذ عمليات بحث عن التشابه عالية الأداء من خلال فهرسة المتجهات وإعادة ترتيب النتائج. كما يتيح إثراء البيانات وتصنيفها والإشراف على المحتوى عن طريق استدعاء نماذج خارجية عبر SQL.

Features

  • AI Model Integrations - Integrates external AI models directly into PostgreSQL via SQL queries for natural language processing and embedding generation.
  • Database AI Toolkits - Provides a comprehensive set of tools to integrate LLMs and vector embeddings directly into PostgreSQL.
  • AI Model Integrations - Integrates external machine learning models directly into PostgreSQL queries for data enrichment and classification.
  • SQL-Based Machine Learning - Enables executing machine learning model requests and inference directly within standard SQL queries.
  • Standard RAG Development - Builds retrieval augmented generation pipelines that combine database retrieval with language models for grounded responses.
  • Retrieval-Augmented Generation - Implements the full retrieval-augmented generation pipeline by combining semantic search results with language model prompts.
  • SQL-Based Model Invocations - Allows executing model requests for text generation, classification, and moderation directly within SQL queries.
  • RAG Pipelines - Implements workflows that augment language model outputs by retrieving and integrating relevant external database data.
  • RAG Application Frameworks - Provides a framework to build RAG applications by combining retrieved database context with model prompts.
  • RAG Frameworks - Provides a framework for building retrieval-augmented generation applications using database context and LLM prompts.
  • SQL-Based Model Invocations - Executes requests to external machine learning models directly from within data queries.
  • Text-to-SQL Translators - Translates natural language user queries into executable SQL statements using semantic catalogs and schema descriptions.
  • Semantic Catalogs - Maintains semantic catalogs that map database objects to natural language descriptions to improve text-to-SQL accuracy.
  • Vector Embeddings - Generates numerical vector representations of database tables and files to enable semantic search.
  • Semantic Vector Search - Retrieves relevant data by calculating the mathematical distance between query and document embeddings.
  • Vector Similarity Search - Performs high-dimensional similarity searches to retrieve data based on semantic meaning rather than keywords.
  • Automatic Vector Embeddings - Automatically converts database content into vector representations using a background worker process.
  • Embedding Synchronization - Automatically synchronizes vector embeddings as source table data changes using state-based tracking.
  • Document Chunking and Embedding Pipelines - Provides automated pipelines that handle the full flow of chunking, embedding, and storing document data.
  • Embedding Pipelines - Implements modular pipelines that automate the loading, parsing, and formatting of data into vector embeddings.
  • Vector-Augmented Queries - Combines semantic vector search with model prompting within standard database queries to build knowledge-aware applications.
  • Semantic Search Engines - Implements a semantic search engine that retrieves information based on conceptual meaning using vector embeddings.
  • Continuous Sync Engines - Implements engines that automatically synchronize vector embeddings as the underlying source data changes.
  • In-Database Model Invocation - Enables executing external machine learning model requests and text-to-SQL translations directly within standard database queries.
  • Vector Indexing - Creates and manages indexes optimized for high-dimensional vector data to support semantic search.
  • Embedding Generation - Provides an automated pipeline to convert database content into high-dimensional vector representations using external workers.
  • Recursive Text Splitting - Implements recursive splitting functions to divide large bodies of text into chunks for AI model consumption.
  • Text Chunks - Splits long text into smaller segments using configurable algorithms and metadata injection for RAG context windows.
  • Result Reranking - Scores and reorders search results against a query to improve precision and relevance.
  • Embedding Status Monitors - Tracks the status of asynchronous embedding jobs, including pending items and processing states.
  • Data Processing - Uses large language models to enrich relational data through automated summarization, categorization, and content moderation.
  • Document Parsing and Extraction - Extracts text from unstructured PDFs and images using OCR to prepare content for vectorization and LLM ingestion.
  • Schema Description Generators - Generates natural language descriptions of database objects to help AI models understand technical schemas.
  • Embedding Synchronization Schedulers - Automates the periodic processing of updated data to synchronize embeddings via a background job system.
  • Document and Unstructured Extraction - Extracts text from PDFs and images using OCR and layout-aware parsing to prepare unstructured content for embedding.
  • Batch Embedding Management - Manages large-scale batch processing of embeddings with built-in resilience against failures and API rate limits.
  • Query Performance Tuning - Optimizes vector query performance by creating specialized indexes on embedding columns to reduce latency.
  • Semantic Knowledge Base Search - Retrieves relevant context and domain knowledge from the database using natural language queries and embeddings.
  • Background Job Processing - Provides a background worker system to process large-scale embedding tasks and synchronization jobs outside the main request flow.
  • Catalog Management - Enables the management of multiple independent semantic catalogs and embedding configurations within one environment.
  • Spatial and Vector Data - Simplifies creating and synchronizing vector embeddings.

سجل النجوم

مخطط تاريخ النجوم لـ timescale/pgaiمخطط تاريخ النجوم لـ timescale/pgai

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ Pgai

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Pgai.
  • casibase/casibaseالصورة الرمزية لـ casibase

    casibase/casibase

    4,443عرض على GitHub↗

    Casibase is an open-source platform that orchestrates multi-turn conversations with large language models and manages retrieval-augmented knowledge bases from a single interface. It provides a unified system for connecting to over 30 AI model providers, ingesting documents into vector embeddings for semantic search, and running autonomous agent loops that can drive a browser, search the web, execute commands, and integrate with external tools. The platform distinguishes itself by combining AI conversation management with infrastructure and application orchestration capabilities. It includes a

    Goa2aagentagi
    عرض على GitHub↗4,443
  • chonkie-inc/chonkieالصورة الرمزية لـ chonkie-inc

    chonkie-inc/chonkie

    4,170عرض على GitHub↗

    Chonkie is a text chunking library designed for retrieval-augmented generation pipelines. It functions as a semantic text splitter and RAG ingestion pipeline, transforming raw text into embedded segments for storage in vector databases. The project distinguishes itself through specialized splitting strategies, including an AST-based code splitter for preserving logical boundaries in source code and a semantic text splitter that uses embedding models to determine boundaries based on meaning. It also provides a vector database ingestor to automate the generation of embeddings and their export t

    Pythonaichonkiechunker
    عرض على GitHub↗4,170
  • docker/genai-stackالصورة الرمزية لـ docker

    docker/genai-stack

    5,333عرض على GitHub↗

    This project is a containerized development stack and application framework for building retrieval-augmented generation systems. It provides a dockerized AI sandbox that integrates local model runtimes, knowledge graphs, and vector stores to enable the creation of contextual chatbots. The stack is distinguished by its graph-based vector store, which combines structured knowledge graphs with vector indices for both semantic and structural data retrieval. It allows for local model hosting with CPU or GPU acceleration, enabling generative tasks without reliance on external cloud APIs. The frame

    Python
    عرض على GitHub↗5,333
  • datawhalechina/all-in-ragالصورة الرمزية لـ datawhalechina

    datawhalechina/all-in-rag

    3,989عرض على GitHub↗

    This project is a retrieval augmented generation framework designed to build pipelines that connect unstructured data and knowledge graphs with large language models. It functions as a vector database orchestrator for indexing text and multimodal content, as well as a system for translating natural language queries into structured database commands. The framework integrates a hybrid retrieval engine that combines dense vector search with sparse keyword matching to increase the precision of retrieved contexts. It further enhances reasoning and relationship mapping through a graph-augmented ret

    Pythonaideepseekembedding
    عرض على GitHub↗3,989
عرض جميع البدائل الـ 30 لـ Pgai→

الأسئلة الشائعة

ما هي وظيفة timescale/pgai؟

pgai هو مجموعة أدوات وإطار عمل لـ PostgreSQL مصمم لدمج نماذج اللغات الكبيرة وتضمينات المتجهات (vector embeddings) مباشرة داخل قاعدة البيانات. يعمل كجسر لتنفيذ طلبات نماذج تعلم الآلة وإجراء ترجمات النص إلى SQL ضمن استعلامات قاعدة البيانات القياسية.

ما هي الميزات الرئيسية لـ timescale/pgai؟

الميزات الرئيسية لـ timescale/pgai هي: AI Model Integrations, Database AI Toolkits, SQL-Based Machine Learning, Standard RAG Development, Retrieval-Augmented Generation, SQL-Based Model Invocations, RAG Pipelines, RAG Application Frameworks.

ما هي البدائل مفتوحة المصدر لـ timescale/pgai؟

تشمل البدائل مفتوحة المصدر لـ timescale/pgai: casibase/casibase — Casibase is an open-source platform that orchestrates multi-turn conversations with large language models and manages… chonkie-inc/chonkie — Chonkie is a text chunking library designed for retrieval-augmented generation pipelines. It functions as a semantic… docker/genai-stack — This project is a containerized development stack and application framework for building retrieval-augmented… datawhalechina/all-in-rag — This project is a retrieval augmented generation framework designed to build pipelines that connect unstructured data… ravendb/ravendb — RavenDB is a multi-model NoSQL document database designed for high-performance, ACID-compliant data storage. It… azure-samples/azure-search-openai-demo — This project is a reference implementation and application template for Retrieval-Augmented Generation (RAG). It…