awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

15 مستودعات

Awesome GitHub RepositoriesMultimodal Models

Models capable of processing and generating across text, image, and other modalities.

Explore 15 awesome GitHub repositories matching part of an awesome list · Multimodal Models. Refine with filters or upvote what's useful.

Awesome Multimodal Models GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • openbmb/minicpm-vالصورة الرمزية لـ OpenBMB

    OpenBMB/MiniCPM-V

    25,653عرض على GitHub↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Efficient multimodal model for visual and textual tasks.

    Python
    عرض على GitHub↗25,653
  • zai-org/open-autoglmالصورة الرمزية لـ zai-org

    zai-org/Open-AutoGLM

    23,532عرض على GitHub↗

    Open-AutoGLM is an autonomous agent framework designed to perform complex user workflows on mobile devices. By translating natural language instructions into precise sequences of taps, scrolls, and text inputs, the system enables the automation of mobile application interactions and testing. The platform distinguishes itself through a combination of vision-language processing and reinforcement learning. It converts graphical user interfaces into structured data, allowing agents to parse screen elements and map natural language commands to coordinate-based actions. To ensure reliability, the s

    Agentic multimodal model for automated device interaction.

    Pythonagentphone-use-agent
    عرض على GitHub↗23,532
  • deepseek-ai/deepseek-ocrالصورة الرمزية لـ deepseek-ai

    deepseek-ai/DeepSeek-OCR

    22,498عرض على GitHub↗

    DeepSeek-OCR is a vision processing framework designed to convert image-based text into machine-readable tokens for large language models. It functions as a document inference pipeline that encodes visual data into compact representations, enabling automated optical character recognition and document analysis workflows. The system distinguishes itself through a high-throughput architecture that utilizes hardware-accelerated batch inference to process large volumes of visual data. It incorporates dynamic resolution scaling to manage the balance between visual detail and token consumption, ensu

    Specialized multimodal model for optical character recognition.

    Python
    عرض على GitHub↗22,498
  • microsoft/unilmالصورة الرمزية لـ microsoft

    microsoft/unilm

    22,030عرض على GitHub↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Framework for transformer-based optical character recognition.

    Pythonbeitbeit-3bitnet
    عرض على GitHub↗22,030
  • qwenlm/qwen2-vlالصورة الرمزية لـ QwenLM

    QwenLM/Qwen2-VL

    19,404عرض على GitHub↗

    Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw

    Multimodal model supporting video and image-text processing.

    Jupyter Notebook
    عرض على GitHub↗19,404
  • bradyfu/awesome-multimodal-large-language-modelsالصورة الرمزية لـ BradyFU

    BradyFU/Awesome-Multimodal-Large-Language-Models

    17,892عرض على GitHub↗

    :sparkles::sparkles:Latest Advances on Multimodal Large Language Models

    Collection of papers and datasets for multimodal language models.

    chain-of-thoughtin-context-learninginstruction-following
    عرض على GitHub↗17,892
  • opengvlab/internvlالصورة الرمزية لـ OpenGVLab

    OpenGVLab/InternVL

    10,061عرض على GitHub↗

    InternVL is a vision-language model framework that fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning. It provides a system for multimodal inference and dialogue, enabling the processing of images and text to answer questions or generate descriptions. The project is distinguished by its high-resolution image processing, which uses dynamic tiling to maintain detail for images up to 4K resolution, and its chain-of-thought visual reasoning for solving complex mathematical and spatial problems. It also supports temporal frame sampling

    Large-scale multimodal model for visual and textual reasoning.

    Pythongptgpt-4ogpt-4v
    عرض على GitHub↗10,061
  • facebookresearch/imagebindالصورة الرمزية لـ facebookresearch

    facebookresearch/ImageBind

    9,036عرض على GitHub↗

    ImageBind is a multi-modal embedding model and joint representation learner that maps images, text, audio, and other modalities into a single shared vector space. It functions as a cross-modal retrieval framework designed to bind multiple sensory inputs into one cohesive mathematical embedding. The system uses a contrastive learning architecture to align disparate data types by maximizing the similarity between related samples. This allows the model to perform zero-shot multimodal classification and execute cross-modal data retrieval, such as locating visual content via natural language descr

    Embedding space model for binding multiple data modalities.

    Python
    عرض على GitHub↗9,036
  • bytedance/dolphinالصورة الرمزية لـ bytedance

    bytedance/Dolphin

    8,820عرض على GitHub↗

    Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital document images into machine-readable structured data. It functions as an LLM document parser, utilizing vision-language models to simultaneously predict spatial layout and text content. The system is designed as a concurrent document processor, employing parallel document parsing to process multiple elements across distributed compute nodes. This high-throughput approach reduces the total time required to convert large volumes of images into structured formats. The project covers

    Multimodal model for image and text integration.

    Pythondocument-analysislayout-analysisocr
    عرض على GitHub↗8,820
  • paddlepaddle/ernieالصورة الرمزية لـ PaddlePaddle

    PaddlePaddle/ERNIE

    7,717عرض على GitHub↗

    ERNIE is a development toolkit for training, fine-tuning, and deploying large language models built on the PaddlePaddle deep learning platform. It provides a comprehensive suite of core components, including an inference server for vision and language models, a training and fine-tuning toolkit, and a framework for building retrieval-augmented generation systems using private knowledge bases. The project features multimodal AI models capable of reasoning across text, images, and video to perform complex visual understanding and information extraction. It distinguishes itself through specialize

    Implements a model architecture capable of reasoning across text, images, and video for visual information extraction.

    Pythonernieernie-45ernie-45-vl
    عرض على GitHub↗7,717
  • qwenlm/qwen-imageالصورة الرمزية لـ QwenLM

    QwenLM/Qwen-Image

    7,379عرض على GitHub↗

    Qwen-Image is a text-to-image model and large language model image generation framework. It functions as an AI image editing suite and a personalized image trainer, capable of producing high-fidelity visuals and accurate typography from natural language descriptions. The system is distinguished by its precision text rendering engine, which integrates multi-script calligraphy and layout-coherent alphabetic text into images. It provides specialized capabilities for subject identity preservation and consistent subject generation across different poses and viewpoints, alongside a training pipelin

    Multimodal model with advanced image understanding capabilities.

    Python
    عرض على GitHub↗7,379
  • facebookresearch/pythiaالصورة الرمزية لـ facebookresearch

    facebookresearch/pythia

    5,635عرض على GitHub↗

    Pythia هو إطار عمل بحثي متعدد الوسائط ونظام تدريب موزع مصمم لبناء وتدريب وتقييم نماذج كبيرة تجمع بين البيانات البصرية واللغوية. يوفر بيئة معيارية لتطوير نماذج الرؤية واللغة، مع التركيز على دمج مدخلات الصور والنصوص في تمثيلات ميزات مشتركة. يستخدم إطار العمل بنية معيارية تفصل كتل بناء النموذج إلى مكونات قابلة للتبديل، مما يسمح بتكوين مرن لوحدات الرؤية واللغة. ويتضمن مجموعة معيارية لتنفيذ النماذج المرجعية مقابل مجموعات بيانات موحدة لإنشاء خطوط أساس أداء متسقة لمهام الرؤية واللغة. يدعم النظام خطوط أنابيب التدريب الموزعة لتوسيع نطاق تطوير النموذج عبر عقد حوسبة متعددة ويستخدم ملفات إعدادات خارجية لتعيين المعلمات الفائقة لضمان قابلية تكرار البحث.

    Provides a modular environment for building and training models capable of processing text and images.

    Python
    عرض على GitHub↗5,635
  • baidubce/qianfan-vlالصورة الرمزية لـ baidubce

    baidubce/Qianfan-VL

    402عرض على GitHub↗

    Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

    Multimodal model specialized for document and visual analysis.

    عرض على GitHub↗402
  • allenai/unified-io-inferenceالصورة الرمزية لـ allenai

    allenai/unified-io-inference

    231عرض على GitHub↗

    This repo contains code to run models from our paper Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

    Unified model for diverse vision and language tasks.

    Jupyter Notebook
    عرض على GitHub↗231
  • tencent/hy-world-2.0T

    Tencent/HY-World-2.0

    0عرض على GitHub↗

    Multimodal model focused on 3D world understanding.

    عرض على GitHub↗0
  1. Home
  2. Part of an Awesome List
  3. AI & Machine Learning
  4. Multimodal Models