17 مستودعات
Configurations for vectorizing data to support semantic search and memory retrieval.
Distinguishing note: Focuses on the configuration of embedding models.
Explore 17 awesome GitHub repositories matching artificial intelligence & ml · Embedding Models. Refine with filters or upvote what's useful.
This repository serves as a comprehensive library of architectural blueprints and code examples for integrating large language models into software applications. It functions as a developer learning resource, providing structured tutorials and implementation patterns that demonstrate how to build intelligent features using advanced prompting and data processing techniques. The collection distinguishes itself by focusing on complex reasoning and data-grounding workflows. It provides practical guidance on implementing retrieval-augmented generation pipelines, which connect language models to pr
Converts unstructured text into numerical representations to enable semantic search and retrieval.
This project is an autonomous agent framework designed to integrate large language models with popular messaging platforms. It functions as a middleware platform that enables automated, multimodal interactions by decomposing complex user goals into sequential plans, executing them through external tools, and maintaining persistent context across sessions. The framework distinguishes itself through a modular skill architecture and a hybrid memory system. Users can extend system capabilities by installing custom logic modules from community hubs or generating them through natural language. The
Agent framework allows specification of provider and dimensions for multimodal embedding models in the configuration file to support memory indexing.
This project is a feature-rich Go client library designed for interacting with Redis. It serves as a comprehensive interface for managing remote data stores, enabling developers to execute standard database commands, handle complex data structures, and perform asynchronous operations within Go applications. The library distinguishes itself through its support for advanced Redis capabilities, including connection pooling, pipelining, and transactional integrity. It provides specialized primitives for managing distributed clusters, including automated topology updates and request routing to sha
Provides configurations for vectorizing data to support semantic search and memory retrieval.
This project is an agentic workflow orchestrator designed for building and deploying autonomous systems that perform multi-step reasoning. It functions as a tool-augmented engine, enabling developers to chain model calls with external function execution to complete complex, user-defined tasks. By integrating large language models with persistent memory and stateful logic, the framework supports the creation of intelligent applications capable of independent operation. The platform distinguishes itself through graph-based state orchestration, which allows developers to define logic steps and t
Maps text, images, and audio into a unified embedding space for advanced semantic search and retrieval systems.
LangBot is an orchestration platform designed for building, managing, and deploying AI agents. It functions as a comprehensive framework for integrating large language models with custom workflows, enabling developers to connect intelligent agents to various messaging platforms and external tools. The platform distinguishes itself through a modular, plugin-based architecture that allows for the extension of agent capabilities via custom tools and file parsers. It features a secure, sandbox-isolated runtime environment that executes untrusted code and plugin logic within resource-constrained c
Provides configurations for vectorizing data to support semantic search and memory retrieval.
LanceDB is a vector database and columnar data store designed to function as a versioned dataset manager and vector search engine. It serves as a high-performance backend for indexing and retrieving high-dimensional embeddings, providing the foundation for machine learning data pipelines. The system distinguishes itself through a combination of cloud-native object storage and immutable version tracking, allowing for data time-travel and reproducible AI experiments. It integrates hybrid search capabilities, merging dense vector similarity with BM25 full-text search and SQL-like scalar filters
Integrates with external embedding model providers to convert raw data into vector representations.
BERTopic is a topic modeling library used to extract interpretable themes from collections of text documents and images. It functions as a document clustering framework that transforms unstructured data into numerical vectors to group semantically similar content. The project distinguishes itself through a multimodal embedding tool that allows for joint clustering of text and images in a shared vector space. It also features a class-based TF-IDF representation engine to identify representative words for clusters and an integrated system for using large language models to generate natural lang
Converts large sentence transformer models into smaller static versions based on specific vocabularies to increase speed.
Torchtune is a PyTorch-native library for fine-tuning, aligning, and quantizing large language models. It provides a config-driven system for instantiating components, orchestrating distributed training, and managing parameter-efficient fine-tuning with quantization support, all through YAML-based configurations and command-line overrides. The library distinguishes itself through its comprehensive post-training workflow orchestration, combining supervised fine-tuning, preference optimization (DPO, PPO, GRPO), knowledge distillation, and quantization-aware training in a single configurable pip
Transfers knowledge from larger teacher models to smaller student models through knowledge distillation.
Torchtune is a PyTorch-native library for fine-tuning, aligning, and quantizing large language models. It provides a configurable training pipeline orchestrated through YAML recipes, with CLI overrides and component swapping, distributed training via FSDP2, memory optimizations, and parameter-efficient fine-tuning methods like LoRA, DoRA, and QLoRA. The library distinguishes itself through its YAML-driven configuration system that defines all training parameters and instantiates components from config files, with full CLI override capability for any field or component at launch time. It suppo
Produces compact student models with competitive performance via teacher-student knowledge transfer.
LLMLingua is a prompt compression tool that reduces token count in prompts before they are sent to a large language model, cutting API costs and latency while preserving task performance. It operates as an extractive pipeline using a BERT-level Transformer encoder to classify each token for removal based on full bidirectional context from the prompt, retaining only key information and discarding non-essential tokens. The tool is trained through a knowledge distillation process, where a compact compression model learns from an extractive dataset derived from a large language model's output to
Trains a compact model on an extractive dataset derived from LLM output to guide token removal.
text2vec is a text vectorization toolkit and semantic similarity framework used to convert words and sentences into numerical vectors. It provides integrated toolsets for generating embeddings, calculating semantic closeness, and implementing lexical and semantic search. The project includes a model fine-tuning pipeline for optimizing embedding and matching models using supervised or unsupervised datasets. It further distinguishes itself by providing a text embedding API that allows vectorization models to be deployed as network services via gRPC or HTTP protocols. The framework covers a bro
Implements model distillation to create smaller, faster versions of embedding models without significant accuracy loss.
fast-reid is a PyTorch-based computer vision framework designed for building, training, and deploying deep learning models for identity-based vision tasks. It provides a specialized toolbox for person re-identification and vehicle re-identification, enabling the matching of individuals and vehicles across non-overlapping camera views. The project includes tools for person attribute recognition to identify specific physical characteristics and traits. It features a modular model zoo that allows for the swapping and benchmarking of different re-identification architectures. The framework cover
Implements model distillation to create efficient versions of complex vision models for faster inference.
nano-graphrag هو نظام استرجاع يستخدم الرسوم البيانية المعرفية لتوفير سياق منظم لاستجابات النماذج اللغوية الكبيرة. يعمل كفهرس للرسوم البيانية المعرفية يحول النص غير المنظم إلى شبكة من الكيانات والعلاقات، بالإضافة إلى نظام استرجاع رسوم بيانية هجين. يتميز المشروع بدمج عمليات البحث في الأحياء المحلية مع ملخصات المجتمع العالمية للإجابة على أسئلة اللغة الطبيعية المعقدة. يتضمن مصوراً للرسوم البيانية المعرفية يولد تمثيلات HTML للكيانات وعلاقاتها لرسم المعرفة المفهرسة. يغطي إطار العمل مجموعة واسعة من الإمكانات بما في ذلك استخراج علاقات الكيانات، وتجميع الرسوم البيانية القائم على المجتمع، والفهرسة التزايدية القائمة على التجزئة. يوفر طبقة تكامل لربط النماذج مفتوحة المصدر وموفري التضمين المحليين، مدعوماً بخلفيات تخزين قابلة للتوصيل لبيانات القيمة المفتاحية، والمتجهات، والرسوم البيانية. يتم توفير فائدة إضافية من خلال التخزين المؤقت للاستجابة القائم على الوسائط ووظائف ما بعد المعالجة لإصلاح مخرجات JSON غير المستقرة من النماذج اللغوية.
Allows swapping between custom language models and embedding providers to control costs or utilize local models.
FastVideo is a comprehensive system for accelerated video generation, serving as a video generation inference engine, a video diffusion training framework, and a modular pipeline orchestrator. It provides a distributed transformer optimizer and a distillation toolkit designed to reduce denoising steps and model complexity to increase frame rates. The project distinguishes itself through specialized acceleration techniques, including joint distillation and sparse attention training. It implements low-step video generation and weight quantization to FP8 or FP4 precision to increase throughput a
Provides tools to reduce denoising steps and model complexity to increase video generation frame rates.
MIRIX is an AI agent state orchestrator and long-term memory system designed to provide persistent context for large language models. It functions as a multi-modal AI memory pipeline that processes text, voice, and screen captures into structured knowledge stores, including a dedicated screen activity knowledge base. The project distinguishes itself by integrating a multi-modal observation pipeline that monitors desktop activity in real-time to build a searchable history of user actions. It utilizes a multi-tiered memory hierarchy—separating episodic, semantic, procedural, and core stores—and
Provides configurations for embedding models to support semantic search and long-term memory retrieval.
Knowledge-Distillation-Zoo هو إطار عمل لضغط نماذج الشبكات العصبية يسهل نقل الأنماط المكتسبة من نماذج المعلم الكبيرة إلى هياكل الطالب الأصغر. يوفر بيئة معيارية لتنفيذ خطوط أنابيب التدريب المصممة لتقليل المتطلبات الحسابية لنماذج التعلم العميق مع الحفاظ على الدقة التنبؤية. تنفذ المكتبة نقل المعرفة من خلال المحاكاة القائمة على logit ومحاذاة خريطة الميزات، مما يسمح للطلاب بتكرار سلوك التصنيف والتمثيلات الداخلية للمعلم. وهي تدعم فصل المعلم عن الطالب، حيث يظل نموذج المعلم مجمداً أثناء عملية التدريب، وتستخدم تكوين الخسارة المعياري لموازنة أهداف المهمة المحددة مع العقوبات الخاصة بالتقطير. تتضمن مجموعة الأدوات واجهة سطر أوامر لإدارة سير عمل التدريب وتدعم حقن المعلمات القابلة للتهيئة للتبديل بين استراتيجيات التقطير المختلفة. تم بناؤها كمكتبة لـ PyTorch، مما يوفر بيئة منظمة لتحسين الشبكات العصبية للنشر على الأجهزة ذات الموارد المحدودة.
Transfers learned patterns from a large teacher model into a smaller student model while maintaining accuracy.
Superlinked is a development framework designed for building semantic search and retrieval pipelines. It functions as a machine learning data pipeline and semantic retrieval engine, providing the tools necessary to unify data schema definition, embedding generation, and vector database integration within a single application. The framework distinguishes itself by acting as a vector database orchestrator that manages the lifecycle of machine learning models alongside complex search logic. It enables developers to construct structured data models that map raw content and metadata into unified r
Handles the lifecycle, warm-up routines, and persistence of embedding models used for generating vector representations.