awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

15 repository-uri

Awesome GitHub RepositoriesMultimodal Models

Models capable of processing and generating across text, image, and other modalities.

Explore 15 awesome GitHub repositories matching part of an awesome list · Multimodal Models. Refine with filters or upvote what's useful.

Awesome Multimodal Models GitHub Repositories

Găsește cele mai bune repo-uri cu AI.Vom căuta cele mai potrivite repository-uri folosind AI.
  • openbmb/minicpm-vAvatar OpenBMB

    OpenBMB/MiniCPM-V

    25,653Vezi pe GitHub↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Efficient multimodal model for visual and textual tasks.

    Python
    Vezi pe GitHub↗25,653
  • zai-org/open-autoglmAvatar zai-org

    zai-org/Open-AutoGLM

    23,532Vezi pe GitHub↗

    Open-AutoGLM is an autonomous agent framework designed to perform complex user workflows on mobile devices. By translating natural language instructions into precise sequences of taps, scrolls, and text inputs, the system enables the automation of mobile application interactions and testing. The platform distinguishes itself through a combination of vision-language processing and reinforcement learning. It converts graphical user interfaces into structured data, allowing agents to parse screen elements and map natural language commands to coordinate-based actions. To ensure reliability, the s

    Agentic multimodal model for automated device interaction.

    Pythonagentphone-use-agent
    Vezi pe GitHub↗23,532
  • deepseek-ai/deepseek-ocrAvatar deepseek-ai

    deepseek-ai/DeepSeek-OCR

    22,498Vezi pe GitHub↗

    DeepSeek-OCR is a vision processing framework designed to convert image-based text into machine-readable tokens for large language models. It functions as a document inference pipeline that encodes visual data into compact representations, enabling automated optical character recognition and document analysis workflows. The system distinguishes itself through a high-throughput architecture that utilizes hardware-accelerated batch inference to process large volumes of visual data. It incorporates dynamic resolution scaling to manage the balance between visual detail and token consumption, ensu

    Specialized multimodal model for optical character recognition.

    Python
    Vezi pe GitHub↗22,498
  • microsoft/unilmAvatar microsoft

    microsoft/unilm

    22,030Vezi pe GitHub↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Framework for transformer-based optical character recognition.

    Pythonbeitbeit-3bitnet
    Vezi pe GitHub↗22,030
  • qwenlm/qwen2-vlAvatar QwenLM

    QwenLM/Qwen2-VL

    19,404Vezi pe GitHub↗

    Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw

    Multimodal model supporting video and image-text processing.

    Jupyter Notebook
    Vezi pe GitHub↗19,404
  • bradyfu/awesome-multimodal-large-language-modelsAvatar BradyFU

    BradyFU/Awesome-Multimodal-Large-Language-Models

    17,892Vezi pe GitHub↗

    :sparkles::sparkles:Latest Advances on Multimodal Large Language Models

    Collection of papers and datasets for multimodal language models.

    chain-of-thoughtin-context-learninginstruction-following
    Vezi pe GitHub↗17,892
  • opengvlab/internvlAvatar OpenGVLab

    OpenGVLab/InternVL

    10,061Vezi pe GitHub↗

    InternVL is a vision-language model framework that fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning. It provides a system for multimodal inference and dialogue, enabling the processing of images and text to answer questions or generate descriptions. The project is distinguished by its high-resolution image processing, which uses dynamic tiling to maintain detail for images up to 4K resolution, and its chain-of-thought visual reasoning for solving complex mathematical and spatial problems. It also supports temporal frame sampling

    Large-scale multimodal model for visual and textual reasoning.

    Pythongptgpt-4ogpt-4v
    Vezi pe GitHub↗10,061
  • facebookresearch/imagebindAvatar facebookresearch

    facebookresearch/ImageBind

    9,036Vezi pe GitHub↗

    ImageBind is a multi-modal embedding model and joint representation learner that maps images, text, audio, and other modalities into a single shared vector space. It functions as a cross-modal retrieval framework designed to bind multiple sensory inputs into one cohesive mathematical embedding. The system uses a contrastive learning architecture to align disparate data types by maximizing the similarity between related samples. This allows the model to perform zero-shot multimodal classification and execute cross-modal data retrieval, such as locating visual content via natural language descr

    Embedding space model for binding multiple data modalities.

    Python
    Vezi pe GitHub↗9,036
  • bytedance/dolphinAvatar bytedance

    bytedance/Dolphin

    8,820Vezi pe GitHub↗

    Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital document images into machine-readable structured data. It functions as an LLM document parser, utilizing vision-language models to simultaneously predict spatial layout and text content. The system is designed as a concurrent document processor, employing parallel document parsing to process multiple elements across distributed compute nodes. This high-throughput approach reduces the total time required to convert large volumes of images into structured formats. The project covers

    Multimodal model for image and text integration.

    Pythondocument-analysislayout-analysisocr
    Vezi pe GitHub↗8,820
  • paddlepaddle/ernieAvatar PaddlePaddle

    PaddlePaddle/ERNIE

    7,717Vezi pe GitHub↗

    ERNIE is a development toolkit for training, fine-tuning, and deploying large language models built on the PaddlePaddle deep learning platform. It provides a comprehensive suite of core components, including an inference server for vision and language models, a training and fine-tuning toolkit, and a framework for building retrieval-augmented generation systems using private knowledge bases. The project features multimodal AI models capable of reasoning across text, images, and video to perform complex visual understanding and information extraction. It distinguishes itself through specialize

    Implements a model architecture capable of reasoning across text, images, and video for visual information extraction.

    Pythonernieernie-45ernie-45-vl
    Vezi pe GitHub↗7,717
  • qwenlm/qwen-imageAvatar QwenLM

    QwenLM/Qwen-Image

    7,379Vezi pe GitHub↗

    Qwen-Image is a text-to-image model and large language model image generation framework. It functions as an AI image editing suite and a personalized image trainer, capable of producing high-fidelity visuals and accurate typography from natural language descriptions. The system is distinguished by its precision text rendering engine, which integrates multi-script calligraphy and layout-coherent alphabetic text into images. It provides specialized capabilities for subject identity preservation and consistent subject generation across different poses and viewpoints, alongside a training pipelin

    Multimodal model with advanced image understanding capabilities.

    Python
    Vezi pe GitHub↗7,379
  • facebookresearch/pythiaAvatar facebookresearch

    facebookresearch/pythia

    5,635Vezi pe GitHub↗

    Pythia este un framework de cercetare multimodală și un sistem de antrenament distribuit conceput pentru construirea, antrenarea și evaluarea modelelor mari care combină date vizuale și lingvistice. Oferă un mediu modular pentru dezvoltarea modelelor vision-language, concentrându-se pe integrarea input-urilor de imagine și text în reprezentări de caracteristici partajate. Framework-ul utilizează o arhitectură modulară care decuplează blocurile de construcție ale modelului în componente interschimbabile, permițând configurarea flexibilă a modulelor de viziune și limbaj. Include o suită de benchmark-uri pentru executarea modelelor de referință pe seturi de date standardizate, pentru a stabili linii de bază de performanță consistente pentru sarcinile vision-language. Sistemul suportă pipeline-uri de antrenament distribuit pentru a scala dezvoltarea modelelor pe mai multe noduri de calcul și utilizează fișiere de configurare externe pentru maparea hiperparametrilor, pentru a asigura reproductibilitatea cercetării.

    Provides a modular environment for building and training models capable of processing text and images.

    Python
    Vezi pe GitHub↗5,635
  • baidubce/qianfan-vlAvatar baidubce

    baidubce/Qianfan-VL

    402Vezi pe GitHub↗

    Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

    Multimodal model specialized for document and visual analysis.

    Vezi pe GitHub↗402
  • allenai/unified-io-inferenceAvatar allenai

    allenai/unified-io-inference

    231Vezi pe GitHub↗

    This repo contains code to run models from our paper Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

    Unified model for diverse vision and language tasks.

    Jupyter Notebook
    Vezi pe GitHub↗231
  • tencent/hy-world-2.0T

    Tencent/HY-World-2.0

    0Vezi pe GitHub↗

    Multimodal model focused on 3D world understanding.

    Vezi pe GitHub↗0
  1. Home
  2. Part of an Awesome List
  3. AI & Machine Learning
  4. Multimodal Models