awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

6 रिपॉजिटरी

Awesome GitHub RepositoriesSemantic Document Retrieval

Fetching relevant documents from an index using vector embeddings for semantic similarity.

Distinct from Vector Document Indexing: Distinct from Vector Document Indexing: focuses on the retrieval step using embeddings, not the indexing workflow.

Explore 6 awesome GitHub repositories matching data & databases · Semantic Document Retrieval. Refine with filters or upvote what's useful.

Awesome Semantic Document Retrieval GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • facebookresearch/llama-recipesfacebookresearch का अवतार

    facebookresearch/llama-recipes

    18,379GitHub पर देखें↗

    This repository is a collection of frameworks and guides for Llama models, functioning as a fine-tuning framework, an inference pipeline, and an AI workflow orchestrator. It provides tools for adapting large language models to specific datasets and domains. The project includes a parameter-efficient fine-tuning toolkit that utilizes techniques like low-rank adaptation to reduce memory and compute requirements. It also serves as an implementation guide for retrieval-augmented generation, combining model inference with external data retrieval to improve response accuracy. The capability surfac

    Queries external databases for relevant text chunks using semantic similarity to ground responses.

    Jupyter Notebook
    GitHub पर देखें↗18,379
  • firebase/genkitfirebase का अवतार

    firebase/genkit

    6,121GitHub पर देखें↗

    Genkit is an open-source framework for building AI-powered applications. It provides a unified interface for connecting to hundreds of generative AI models from multiple providers, enabling text, image, audio, and video generation through a single API. The framework structures multi-step AI interactions—including chat, retrieval-augmented generation, tool use, and agentic workflows—as composable, traceable flows with built-in streaming and state management. The framework distinguishes itself through a comprehensive developer toolkit that includes a command-line interface and a local developer

    Fetches relevant documents from an index using vector embeddings for semantic similarity.

    TypeScript
    GitHub पर देखें↗6,121
  • brianpetro/obsidian-smart-connectionsbrianpetro का अवतार

    brianpetro/obsidian-smart-connections

    5,195GitHub पर देखें↗

    यह प्रोजेक्ट एक नॉलेज बेस प्लगइन और RAG कॉन्टेक्स्ट मैनेजर है जो सिमेंटिक सर्च और रिलेशनशिप मैपिंग को सक्षम करने के लिए एक स्थानीय वेक्टर डेटाबेस इंटरफेस का उपयोग करता है। यह कीवर्ड मिलान के बजाय वैचारिक अर्थ के आधार पर सिमेंटिक रूप से संबंधित नोट्स और अंशों को खोजने के लिए टेक्स्ट को संख्यात्मक वैक्टर में बदलता है। यह सिस्टम एक सिमेंटिक ग्राफ़ विज़ुअलाइज़र के माध्यम से खुद को अलग करता है जो वैचारिक कनेक्शन को प्रकट करने के लिए नोट्स को क्लस्टर में मैप करता है। इसमें एक कॉन्टेक्स्ट मैनेजर भी है जो बड़े भाषा मॉडल वार्तालापों के लिए आधारभूत तथ्यात्मक आधार प्रदान करने के लिए स्थानीय नोट्स और अंशों को पुन: प्रयोज्य पैक में बंडल करने में सक्षम है। यह टूल प्राकृतिक भाषा ज्ञान क्वेरी, नोट निर्माण के लिए स्वचालित वर्कफ़्लो निष्पादन, और स्थानीय तथा क्लाउड-आधारित AI मॉडल के बीच प्रॉम्प्ट को रूट करने की क्षमता सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह संपादन प्रक्रिया के दौरान समान दस्तावेज़ों को सतह पर लाने के लिए इनलाइन संबंधित सामग्री संकेतक और एक फ़ूटर पैनल जैसे कई खोज इंटरफेस प्रदान करता है।

    Surfaces semantically similar excerpts based on the active document to discover relevant prior work.

    JavaScriptchatgptclaudeembeddings
    GitHub पर देखें↗5,195
  • facebookresearch/drqafacebookresearch का अवतार

    facebookresearch/DrQA

    4,468GitHub पर देखें↗

    DrQA एक ओपन-डोमेन प्रश्न उत्तर प्रणाली है जो एक बड़े कॉर्पस से प्रासंगिक दस्तावेज़ों को पुनः प्राप्त करती है और प्राकृतिक भाषा के प्रश्नों के विशिष्ट उत्तर निकालती है। इसे एक न्यूरल नेटवर्क सिस्टम के रूप में लागू किया गया है जो मशीन रीडिंग कॉम्प्रिहेंशन मॉडल के साथ एक दस्तावेज़ रिट्रीवल इंजन को जोड़ता है। यह सिस्टम टू-स्टेज पाइपलाइन आर्किटेक्चर का उपयोग करता है। एक कोर्स-ग्रेन्ड दस्तावेज़ रिट्रीवर संभावित दस्तावेज़ों की पहचान करने के लिए वेटेड वर्ड वेक्टर्स का उपयोग करता है, जबकि एक फाइन-ग्रेन्ड मशीन रीडिंग कॉम्प्रिहेंशन मॉडल उत्तर वाले सटीक टेक्स्ट स्पैन की पहचान करता है और उसे निकालता है। इस प्रोजेक्ट में एक सुपरवाइज्ड नेचुरल लैंग्वेज प्रोसेसिंग डेटासेट जनरेटर भी शामिल है। यह टूल स्वचालित स्ट्रिंग ह्यूरिस्टिक्स का उपयोग करके प्रश्न-उत्तर जोड़ियों को सहायक पैराग्राफ के साथ मिलान करके ट्रेनिंग उदाहरण बनाता है। कोडबेस न्यूरल नेटवर्क प्रोसेसिंग के लिए रॉ टेक्स्ट तैयार करने के लिए टेक्स्ट प्रोसेसिंग और टोकनाइज़ेशन के लिए अतिरिक्त क्षमताएं प्रदान करता है।

    Uses vector embeddings for semantic document retrieval within a large unstructured corpus.

    Python
    GitHub पर देखें↗4,468
  • rag-web-ui/rag-web-uirag-web-ui का अवतार

    rag-web-ui/rag-web-ui

    3,048GitHub पर देखें↗

    This project is a web-based platform designed for retrieval-augmented generation, providing a conversational interface that connects language models to private document collections. It functions as a comprehensive system for managing enterprise knowledge bases, allowing users to query internal documents and receive context-aware, source-verified answers through a natural language chat interface. The platform distinguishes itself by integrating vector-based semantic search with modular knowledge base management, enabling the ingestion, segmentation, and indexing of documents into searchable em

    Uses vector embeddings to perform rapid semantic similarity searches during query generation.

    TypeScriptaideepseeklangchain
    GitHub पर देखें↗3,048
  • kennethleungty/llama-2-open-source-llm-cpu-inferencekennethleungty का अवतार

    kennethleungty/Llama-2-Open-Source-LLM-CPU-Inference

    973GitHub पर देखें↗

    यह प्रोजेक्ट स्थानीय उपभोक्ता हार्डवेयर पर पूरी तरह से लार्ज लैंग्वेज मॉडल (LLM) चलाने और दस्तावेज़-आधारित प्रश्न-उत्तर (QA) करने के लिए एक फ्रेमवर्क प्रदान करता है। CPU-आधारित इन्फरेंस इंजन को एक स्थानीय वेक्टर डेटाबेस के साथ जोड़कर, यह उपयोगकर्ताओं को क्लाउड-आधारित API या विशेष GPU पर निर्भर रहे बिना जानकारी प्रोसेस करने की सुविधा देता है। यह सिस्टम एक कमांड-लाइन टूल के रूप में कार्य करता है जो निजी जानकारी प्रोसेसिंग के पूरे लाइफसाइकिल को मैनेज करता है। यह स्थानीय टेक्स्ट फ़ाइलों को खोजने योग्य वेक्टर एम्बेडिंग में बदल देता है, जिससे मॉडल प्रासंगिक संदर्भ (context) प्राप्त कर सकता है और उपयोगकर्ता द्वारा प्रदान की गई सामग्री के आधार पर सटीक उत्तर दे सकता है। क्वांटाइज्ड मॉडल निष्पादन का उपयोग करके, यह फ्रेमवर्क मेमोरी और कंप्यूट की ज़रूरतों को कम करता है ताकि इसे सामान्य हार्डवेयर पर आसानी से चलाया जा सके। इस प्रोजेक्ट में दस्तावेज़ इंडेक्सिंग, सिमेंटिक रिट्रीवल और संदर्भ-जागरूक जनरेशन के लिए एक पूर्ण पाइपलाइन शामिल है। यह सभी दस्तावेज़ इनजेशन, एम्बेडिंग जनरेशन और मॉडल इन्फरेंस कार्यों को स्थानीय वातावरण में रखकर डेटा गोपनीयता सुनिश्चित करता है।

    Fetches relevant document segments from an index using vector embeddings for semantic similarity.

    Pythonc-transformerschatgptcpu
    GitHub पर देखें↗973
  1. Home
  2. Data & Databases
  3. Database Management Systems
  4. Database Engines
  5. Vector Databases
  6. Vector Document Indexing
  7. Semantic Document Retrieval