10 مستودعات
Neural networks that map multiple data modalities into a single shared vector space.
Distinct from Multi-Modal Tokenizers: Focuses on the complete embedding model rather than just the tokenization process
Explore 10 awesome GitHub repositories matching artificial intelligence & ml · Multi-Modal Embedding Models. Refine with filters or upvote what's useful.
ImageBind is a multi-modal embedding model and joint representation learner that maps images, text, audio, and other modalities into a single shared vector space. It functions as a cross-modal retrieval framework designed to bind multiple sensory inputs into one cohesive mathematical embedding. The system uses a contrastive learning architecture to align disparate data types by maximizing the similarity between related samples. This allows the model to perform zero-shot multimodal classification and execute cross-modal data retrieval, such as locating visual content via natural language descr
Maps images, text, audio, and other modalities into a single shared vector space using a neural network.
This repository provides a collection of reference implementations and code examples for training and deploying machine learning models using the MLX framework. It serves as a practical guide for executing distributed training, fine-tuning large language models, converting model weights, and implementing multimodal generative workflows. The project distinguishes itself through specialized examples for local hardware execution, featuring weight quantization to reduce memory usage and low-rank adaptation for parameter-efficient fine-tuning. It also includes scripts for transforming external mod
Implements neural networks that map images and text into a shared vector space for joint retrieval.
lmdeploy is a high-performance inference engine and deployment framework for large language models and vision models. It functions as a multi-modal model server and compression toolkit designed to serve models with high throughput and low latency. The system enables the distribution of model services across multiple machines using request-based load balancing and tensor parallelism. It includes specialized tools for model quantization and compression to reduce the memory footprint of weights and caches. The framework covers broad capability areas including production deployment, distributed
Coordinates the flow of image and text data through distinct encoders before processing them in a unified transformer.
GLM-4 is a large language model and fine-tuning framework designed for human-like text production, complex reasoning, and multilingual conversation. It functions as a multimodal system capable of processing high-resolution visual content and as a long-context model designed to analyze documents with a context window of up to one million tokens. The project differentiates itself through a function calling interface that enables AI agent development by connecting the model to external APIs and real-time web browsing. It includes specialized capabilities for generating functional programming cod
Integrates high-resolution visual features into a shared vector space for joint visual and linguistic reasoning.
Gemma هي عائلة من نماذج اللغات الكبيرة ذات الأوزان المفتوحة القائمة على بنية محول (transformer) فك التشفير فقط. تم تصميم هذه النماذج لتوليد النصوص والمحادثات متعددة الوسائط، وهي قادرة على معالجة وتوليد الردود بناءً على تسلسلات المدخلات النصية والبصرية. يوفر المشروع نموذج ذكاء اصطناعي قابلاً للضبط الدقيق يدعم تعديل الأوزان والتكيف منخفض الرتبة (low-rank adaptation) لتخصيص الأداء لمهام معينة. يتضمن دعماً للأوزان المكممة (quantized) لتقليل استخدام الذاكرة وزيادة سرعة الاستنتاج على الأجهزة المحدودة. تغطي مساحة القدرات تكامل الذكاء الاصطناعي متعدد الوسائط، وتحسين الذاكرة من خلال تجزئة المعلمات، ودمج الأدوات الخارجية وواجهات برمجة التطبيقات لاسترجاع البيانات في الوقت الفعلي. كما تتيح توليد الصور من النصوص وأخذ عينات من مخرجات النصوص المهيكلة.
Integrates visual and textual data by mapping different input modalities into a shared latent space for joint processing.
DeepSeek-VL2 هو نموذج لغوي كبير متعدد الوسائط ونظام رؤية لغوية مصمم لتحليل المشاهد المرئية وتوليد نص وصفي. يعمل كنموذج للإجابة على الأسئلة المرئية والتأريض المرئي، وقادر على استخراج المعلومات من المستندات وتحديد كائنات أو مناطق محددة داخل الصور بناءً على أوصاف نصية. يستخدم المشروع معمارية خليط من الخبراء (mixture-of-experts) لمعالجة مدخلات الصور والنصوص المدمجة. تم تحسينه للاستدلال من خلال التعبئة التزايدية (incremental prefilling)، مما يقلل من متطلبات ذاكرة GPU على الأجهزة. يغطي النموذج تحليل البيانات متعدد الوسائط وفهم المستندات المرئية، بما في ذلك تفسير المخططات والتخطيطات. يقوم بإجراء استدلال مرئي وتأريض لمطابقة الاستعلامات النصية مع المحتوى المرئي المقابل.
Implements cross-attention mechanisms to align visual regions with specific text tokens.
LightGlue هو إطار عمل للتعلم العميق مصمم لمطابقة الميزات المحلية وتقدير المراسلات عالية السرعة بين أزواج الصور. يعمل كنموذج مطابقة للرؤية الحاسوبية يحدد النقاط الرئيسية المتقابلة عبر وجهات نظر مختلفة. يستخدم النظام بنية شبكة عصبية تكيفية تعمل على تحسين سرعة الاستدلال ديناميكيًا عن طريق تقليم عمقها وعرضها بناءً على أزواج الصور المدخلة. يستخدم هذا النهج آلية انتباه بنمط transformer وانتباه عبر الصور لحساب الارتباطات بين واصفات الميزات. تتضمن عملية المطابقة حلقة تحسين تكرارية وإيقافًا مبكرًا ديناميكيًا لإيقاف الحساب بمجرد استيفاء عتبات الثقة. تدعم هذه القدرات خط أنابيب رؤية حاسوبية أوسع لمحاذاة الصور في الوقت الفعلي وتحسين استدلال الشبكة العصبية.
Employs cross-attention mechanisms to compute correlations between feature descriptors of two different images.
sam-hq هي مجموعة من نماذج أساس الرؤية المدربة مسبقاً والمحولات المصممة لتجزئة الصور عالية الجودة، واستخراج الميزات متعددة الوسائط، وتقدير العمق. توفر نموذج رؤية صفري (zero-shot) قادراً على إجراء التجزئة والتصنيف عبر مجالات متنوعة دون الحاجة إلى تدريب خاص بالمهمة. يتميز المشروع بأداة تجزئة صور عالية الجودة تعتمد على Segment Anything Model التي تولد أقنعة دقيقة من توجيهات مكانية. يتضمن مستخرج ميزات متعدد الوسائط لتوليد تضمينات متجهة عالية الأبعاد من مدخلات الصور والنصوص، بالإضافة إلى أداة التفافية للتنبؤ بالمسافة أو ارتفاع المظلة من البيانات البصرية. يغطي إطار العمل مجموعة واسعة من قدرات الرؤية الحاسوبية، بما في ذلك تصنيف الصور، واستخراج الميزات متعددة الدقة، ومعالجة الصور مسبقاً. يدعم التكيف مع المجال من خلال الضبط الدقيق على مجموعات بيانات مخصصة للتطبيقات المتخصصة مثل التصوير الطبي والاستشعار عن بُعد. يمكن تحويل فك تشفير القناع إلى تنسيق مفتوح للتنفيذ في بيئات ذات بيئة تشغيل قياسية.
Provides a multimodal embedding model that maps image and text data into a shared vector space.
VisualGLM-6B هو نموذج لغوي كبير متعدد الوسائط ثنائي اللغة ونموذج رؤية-لغة مصمم لمهام المحادثة والفهم البصري. يعمل كنموذج ذكاء اصطناعي ثنائي اللغة قادر على معالجة وتوليد الاستجابات باللغتين الصينية والإنجليزية. النظام عبارة عن نموذج لغوي كبير مكمم (quantized) يدعم دقة 4-بت و8-بت لتقليل استخدام الذاكرة ومتطلبات الأجهزة أثناء النشر المحلي. وهو أيضاً نموذج فعال في ضبط المعلمات (parameter-efficient fine-tuning)، مما يسمح بتعديلات الأوزان لتكييف النظام مع مهام محددة دون الحاجة لإعادة التدريب الكامل. يغطي المشروع الذكاء الاصطناعي للمحادثة متعدد الوسائط والحوار القائم على الصور، مما يتيح تحليل المحتوى البصري لأداء مهام الفهم البصري عبر لغات متعددة. تشمل قدراته تكميم دقة النموذج والضبط الدقيق الخاص بالمجال للتطبيقات المتخصصة.
Provides a mechanism to transform visual tokens into a sequence the language model can process as words.
Multimodal is a machine learning library built on PyTorch for training large-scale models that combine text, image, audio, and video data streams. It functions as a deep learning framework dedicated to generative diffusion models, multi-task training, and vision-language tasks. The library supplies modular building blocks, discrete latent codebook quantization, shared-space embeddings, and stackable adapter layers to handle diverse conditional inputs during training and inference. The framework supports specific architectures for diffusion models, text-to-video generation, image-text retrieva
Projects different data modalities into a joint vector space to compute similarity scores for cross-modal retrieval tasks.