awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

10 مستودعات

Awesome GitHub RepositoriesMulti-Modal Embedding Models

Neural networks that map multiple data modalities into a single shared vector space.

Distinct from Multi-Modal Tokenizers: Focuses on the complete embedding model rather than just the tokenization process

Explore 10 awesome GitHub repositories matching artificial intelligence & ml · Multi-Modal Embedding Models. Refine with filters or upvote what's useful.

Awesome Multi-Modal Embedding Models GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • facebookresearch/imagebindالصورة الرمزية لـ facebookresearch

    facebookresearch/ImageBind

    9,036عرض على GitHub↗

    ImageBind is a multi-modal embedding model and joint representation learner that maps images, text, audio, and other modalities into a single shared vector space. It functions as a cross-modal retrieval framework designed to bind multiple sensory inputs into one cohesive mathematical embedding. The system uses a contrastive learning architecture to align disparate data types by maximizing the similarity between related samples. This allows the model to perform zero-shot multimodal classification and execute cross-modal data retrieval, such as locating visual content via natural language descr

    Maps images, text, audio, and other modalities into a single shared vector space using a neural network.

    Python
    عرض على GitHub↗9,036
  • ml-explore/mlx-examplesالصورة الرمزية لـ ml-explore

    ml-explore/mlx-examples

    8,254عرض على GitHub↗

    This repository provides a collection of reference implementations and code examples for training and deploying machine learning models using the MLX framework. It serves as a practical guide for executing distributed training, fine-tuning large language models, converting model weights, and implementing multimodal generative workflows. The project distinguishes itself through specialized examples for local hardware execution, featuring weight quantization to reduce memory usage and low-rank adaptation for parameter-efficient fine-tuning. It also includes scripts for transforming external mod

    Implements neural networks that map images and text into a shared vector space for joint retrieval.

    Pythonmlx
    عرض على GitHub↗8,254
  • internlm/lmdeployالصورة الرمزية لـ InternLM

    InternLM/lmdeploy

    7,903عرض على GitHub↗

    lmdeploy is a high-performance inference engine and deployment framework for large language models and vision models. It functions as a multi-modal model server and compression toolkit designed to serve models with high throughput and low latency. The system enables the distribution of model services across multiple machines using request-based load balancing and tensor parallelism. It includes specialized tools for model quantization and compression to reduce the memory footprint of weights and caches. The framework covers broad capability areas including production deployment, distributed

    Coordinates the flow of image and text data through distinct encoders before processing them in a unified transformer.

    Pythoncodellamacuda-kernelsdeepspeed
    عرض على GitHub↗7,903
  • zai-org/glm-4الصورة الرمزية لـ zai-org

    zai-org/GLM-4

    7,058عرض على GitHub↗

    GLM-4 is a large language model and fine-tuning framework designed for human-like text production, complex reasoning, and multilingual conversation. It functions as a multimodal system capable of processing high-resolution visual content and as a long-context model designed to analyze documents with a context window of up to one million tokens. The project differentiates itself through a function calling interface that enables AI agent development by connecting the model to external APIs and real-time web browsing. It includes specialized capabilities for generating functional programming cod

    Integrates high-resolution visual features into a shared vector space for joint visual and linguistic reasoning.

    Pythonchatglmchatglm-6bglm
    عرض على GitHub↗7,058
  • google-deepmind/gemmaالصورة الرمزية لـ google-deepmind

    google-deepmind/gemma

    5,475عرض على GitHub↗

    Gemma هي عائلة من نماذج اللغات الكبيرة ذات الأوزان المفتوحة القائمة على بنية محول (transformer) فك التشفير فقط. تم تصميم هذه النماذج لتوليد النصوص والمحادثات متعددة الوسائط، وهي قادرة على معالجة وتوليد الردود بناءً على تسلسلات المدخلات النصية والبصرية. يوفر المشروع نموذج ذكاء اصطناعي قابلاً للضبط الدقيق يدعم تعديل الأوزان والتكيف منخفض الرتبة (low-rank adaptation) لتخصيص الأداء لمهام معينة. يتضمن دعماً للأوزان المكممة (quantized) لتقليل استخدام الذاكرة وزيادة سرعة الاستنتاج على الأجهزة المحدودة. تغطي مساحة القدرات تكامل الذكاء الاصطناعي متعدد الوسائط، وتحسين الذاكرة من خلال تجزئة المعلمات، ودمج الأدوات الخارجية وواجهات برمجة التطبيقات لاسترجاع البيانات في الوقت الفعلي. كما تتيح توليد الصور من النصوص وأخذ عينات من مخرجات النصوص المهيكلة.

    Integrates visual and textual data by mapping different input modalities into a shared latent space for joint processing.

    Python
    عرض على GitHub↗5,475
  • deepseek-ai/deepseek-vl2الصورة الرمزية لـ deepseek-ai

    deepseek-ai/DeepSeek-VL2

    5,302عرض على GitHub↗

    DeepSeek-VL2 هو نموذج لغوي كبير متعدد الوسائط ونظام رؤية لغوية مصمم لتحليل المشاهد المرئية وتوليد نص وصفي. يعمل كنموذج للإجابة على الأسئلة المرئية والتأريض المرئي، وقادر على استخراج المعلومات من المستندات وتحديد كائنات أو مناطق محددة داخل الصور بناءً على أوصاف نصية. يستخدم المشروع معمارية خليط من الخبراء (mixture-of-experts) لمعالجة مدخلات الصور والنصوص المدمجة. تم تحسينه للاستدلال من خلال التعبئة التزايدية (incremental prefilling)، مما يقلل من متطلبات ذاكرة GPU على الأجهزة. يغطي النموذج تحليل البيانات متعدد الوسائط وفهم المستندات المرئية، بما في ذلك تفسير المخططات والتخطيطات. يقوم بإجراء استدلال مرئي وتأريض لمطابقة الاستعلامات النصية مع المحتوى المرئي المقابل.

    Implements cross-attention mechanisms to align visual regions with specific text tokens.

    Python
    عرض على GitHub↗5,302
  • cvg/lightglueالصورة الرمزية لـ cvg

    cvg/LightGlue

    4,625عرض على GitHub↗

    LightGlue هو إطار عمل للتعلم العميق مصمم لمطابقة الميزات المحلية وتقدير المراسلات عالية السرعة بين أزواج الصور. يعمل كنموذج مطابقة للرؤية الحاسوبية يحدد النقاط الرئيسية المتقابلة عبر وجهات نظر مختلفة. يستخدم النظام بنية شبكة عصبية تكيفية تعمل على تحسين سرعة الاستدلال ديناميكيًا عن طريق تقليم عمقها وعرضها بناءً على أزواج الصور المدخلة. يستخدم هذا النهج آلية انتباه بنمط transformer وانتباه عبر الصور لحساب الارتباطات بين واصفات الميزات. تتضمن عملية المطابقة حلقة تحسين تكرارية وإيقافًا مبكرًا ديناميكيًا لإيقاف الحساب بمجرد استيفاء عتبات الثقة. تدعم هذه القدرات خط أنابيب رؤية حاسوبية أوسع لمحاذاة الصور في الوقت الفعلي وتحسين استدلال الشبكة العصبية.

    Employs cross-attention mechanisms to compute correlations between feature descriptors of two different images.

    Python
    عرض على GitHub↗4,625
  • syscv/sam-hqالصورة الرمزية لـ SysCV

    SysCV/sam-hq

    4,234عرض على GitHub↗

    sam-hq هي مجموعة من نماذج أساس الرؤية المدربة مسبقاً والمحولات المصممة لتجزئة الصور عالية الجودة، واستخراج الميزات متعددة الوسائط، وتقدير العمق. توفر نموذج رؤية صفري (zero-shot) قادراً على إجراء التجزئة والتصنيف عبر مجالات متنوعة دون الحاجة إلى تدريب خاص بالمهمة. يتميز المشروع بأداة تجزئة صور عالية الجودة تعتمد على Segment Anything Model التي تولد أقنعة دقيقة من توجيهات مكانية. يتضمن مستخرج ميزات متعدد الوسائط لتوليد تضمينات متجهة عالية الأبعاد من مدخلات الصور والنصوص، بالإضافة إلى أداة التفافية للتنبؤ بالمسافة أو ارتفاع المظلة من البيانات البصرية. يغطي إطار العمل مجموعة واسعة من قدرات الرؤية الحاسوبية، بما في ذلك تصنيف الصور، واستخراج الميزات متعددة الدقة، ومعالجة الصور مسبقاً. يدعم التكيف مع المجال من خلال الضبط الدقيق على مجموعات بيانات مخصصة للتطبيقات المتخصصة مثل التصوير الطبي والاستشعار عن بُعد. يمكن تحويل فك تشفير القناع إلى تنسيق مفتوح للتنفيذ في بيئات ذات بيئة تشغيل قياسية.

    Provides a multimodal embedding model that maps image and text data into a shared vector space.

    Jupyter Notebookhigh-qualitysamsegment-anything
    عرض على GitHub↗4,234
  • thudm/visualglm-6bالصورة الرمزية لـ THUDM

    THUDM/VisualGLM-6B

    4,157عرض على GitHub↗

    VisualGLM-6B هو نموذج لغوي كبير متعدد الوسائط ثنائي اللغة ونموذج رؤية-لغة مصمم لمهام المحادثة والفهم البصري. يعمل كنموذج ذكاء اصطناعي ثنائي اللغة قادر على معالجة وتوليد الاستجابات باللغتين الصينية والإنجليزية. النظام عبارة عن نموذج لغوي كبير مكمم (quantized) يدعم دقة 4-بت و8-بت لتقليل استخدام الذاكرة ومتطلبات الأجهزة أثناء النشر المحلي. وهو أيضاً نموذج فعال في ضبط المعلمات (parameter-efficient fine-tuning)، مما يسمح بتعديلات الأوزان لتكييف النظام مع مهام محددة دون الحاجة لإعادة التدريب الكامل. يغطي المشروع الذكاء الاصطناعي للمحادثة متعدد الوسائط والحوار القائم على الصور، مما يتيح تحليل المحتوى البصري لأداء مهام الفهم البصري عبر لغات متعددة. تشمل قدراته تكميم دقة النموذج والضبط الدقيق الخاص بالمجال للتطبيقات المتخصصة.

    Provides a mechanism to transform visual tokens into a sequence the language model can process as words.

    Python
    عرض على GitHub↗4,157
  • facebookresearch/multimodalالصورة الرمزية لـ facebookresearch

    facebookresearch/multimodal

    1,723عرض على GitHub↗

    Multimodal is a machine learning library built on PyTorch for training large-scale models that combine text, image, audio, and video data streams. It functions as a deep learning framework dedicated to generative diffusion models, multi-task training, and vision-language tasks. The library supplies modular building blocks, discrete latent codebook quantization, shared-space embeddings, and stackable adapter layers to handle diverse conditional inputs during training and inference. The framework supports specific architectures for diffusion models, text-to-video generation, image-text retrieva

    Projects different data modalities into a joint vector space to compute similarity scores for cross-modal retrieval tasks.

    Python
    عرض على GitHub↗1,723
  1. Home
  2. Artificial Intelligence & ML
  3. Multi-Modal Tokenizers
  4. Multi-Modal Embedding Models

استكشف الوسوم الفرعية

  • Cross-Attention Mechanisms1 وسم فرعيNeural network layers that map different input modalities into a shared latent space for joint processing. **Distinct from Multi-Modal Embedding Models:** Focuses on the attention mechanism that integrates modalities, whereas Multi-Modal Embedding Models refers to the overall model architecture.
  • Multimodal Pipeline CoordinatorsSystems that coordinate the flow of data from multiple encoders into a unified transformer model. **Distinct from Multi-Modal Embedding Models:** Distinct from Multi-Modal Embedding Models: focuses on the operational pipeline and coordination of encoders rather than the embedding model architecture itself.