awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

16 مستودعات

Awesome GitHub RepositoriesMultimodal Diffusion Models

Models integrating vision, language, and action through diffusion processes.

Explore 16 awesome GitHub repositories matching part of an awesome list · Multimodal Diffusion Models. Refine with filters or upvote what's useful.

Awesome Multimodal Diffusion Models GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • vectorspacelab/omnigenالصورة الرمزية لـ VectorSpaceLab

    VectorSpaceLab/OmniGen

    4,326عرض على GitHub↗

    OmniGen is a unified image generation model and diffusion framework that processes text, images, and vision tasks through a single system. It functions as a multimodal diffusion framework that treats diverse vision operations as unified image synthesis problems using shared model weights, removing the need for external adapter modules. The system supports subject-driven image generation to preserve the identity of objects from reference photos and allows for multi-reference image synthesis. It also operates as an instruction-based image editor, modifying visual content through natural languag

    Provides a multimodal diffusion framework integrating vision and language via shared model weights.

    Jupyter Notebookdiffusionimageimage-edit
    عرض على GitHub↗4,326
  • ml-gsai/lladaالصورة الرمزية لـ ML-GSAI

    ML-GSAI/LLaDA

    3,580عرض على GitHub↗

    LLaDA is a masked diffusion language model and conditional text generator. It generates text by iteratively refining masked tokens through a diffusion process rather than predicting the next token in a sequence. The project functions as a vision-language diffusion model, converting visual inputs into text responses. It also serves as a preference optimization framework that uses log-likelihood estimation and evidence lower bounds to tune model responses. The system supports multi-round conversational AI and text sequence evaluation. It integrates vision-language embedding for cross-modal con

    A system that converts visual inputs into text responses using a diffusion process for complex multimodal tasks.

    Python
    عرض على GitHub↗3,580
  • gen-verse/mmadaالصورة الرمزية لـ Gen-Verse

    Gen-Verse/MMaDA

    1,656عرض على GitHub↗

    Multimodal Large Diffusion Language Models (NeurIPS 2025)

    Multimodal large diffusion language model architecture.

    Python
    عرض على GitHub↗1,656
  • alpha-vllm/lumina-dimooالصورة الرمزية لـ Alpha-VLLM

    Alpha-VLLM/Lumina-DiMOO

    1,001عرض على GitHub↗

    Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

    Omni-diffusion model for multimodal generation and understanding.

    Python
    عرض على GitHub↗1,001
  • ml-gsai/llada-vالصورة الرمزية لـ ML-GSAI

    ML-GSAI/LLaDA-V

    345عرض على GitHub↗

    2026.03.23 We are excited to introduce LLaDA-o, the latest model in the LLaDA series. As an effective and length-adaptive omni diffusion model for unified multimodal understanding and generation, LLaDA-o extends the LLaDA line to broader multimodal settings, supporting visual understanding,…

    Large language diffusion models with visual instruction tuning.

    Python
    عرض على GitHub↗345
  • tyfeld/mmada-parallelالصورة الرمزية لـ tyfeld

    tyfeld/MMaDA-Parallel

    300عرض على GitHub↗

    ICLR 2026 MMaDA-Parallel: Parallel Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation

    Multimodal diffusion model for thinking-aware editing and generation.

    Python
    عرض على GitHub↗300
  • jacklishufan/lavidaالصورة الرمزية لـ jacklishufan

    jacklishufan/LaViDa

    221عرض على GitHub↗

    [Paper](paper/paper.pdf) [Arxiv](https://arxiv.org/abs/2505.16839) [Checkpoints](https://huggingface.co/collections/jacklishufan/lavida-10-682ecf5a5fa8c5df85c61ded) [Data](https://huggingface.co/datasets/jacklishufan/lavida-train) [Website](https://homepage.jackli.org/projects/lavida/)

    Large diffusion language model for multimodal understanding.

    Python
    عرض على GitHub↗221
  • openhelix-team/unified-diffusion-vlaالصورة الرمزية لـ OpenHelix-Team

    OpenHelix-Team/Unified-Diffusion-VLA

    182عرض على GitHub↗

    Jiayi Chen¹\,Wenxuan Song¹†\, Pengxiang Ding²˒³, Ziyang Zhou¹, Han Zhao²˒³, Feilong Tang⁴,Donglin Wang², Haoang Li¹‡

    Joint discrete denoising for vision-language-action models.

    Python
    عرض على GitHub↗182
  • hustvl/diffusionvlالصورة الرمزية لـ hustvl

    hustvl/DiffusionVL

    149عرض على GitHub↗

    DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models

    Translating autoregressive models into vision-language diffusion models.

    Python
    عرض على GitHub↗149
  • alexanderswerdlow/unidiscالصورة الرمزية لـ alexanderswerdlow

    alexanderswerdlow/unidisc

    141عرض على GitHub↗

    Unified Multimodal Discrete Diffusion

    Unified multimodal discrete diffusion framework.

    Python
    عرض على GitHub↗141
  • yu-rp/dimpleالصورة الرمزية لـ yu-rp

    yu-rp/Dimple

    117عرض على GitHub↗

    Dimple, the first Discrete Diffusion Multimodal Large Language Model

    Discrete diffusion multimodal model with parallel decoding.

    Python
    عرض على GitHub↗117
  • m-e-agi-lab/mudditالصورة الرمزية لـ M-E-AGI-Lab

    M-E-AGI-Lab/Muddit

    117عرض على GitHub↗

    Muddit is the 2nd generation Meissonic. It is built upon discrete diffusion for unified and efficient multimodal generation.

    Unified discrete diffusion for multimodal generation beyond text-to-image.

    Python
    عرض على GitHub↗117
  • fudoki-hku/fudokiالصورة الرمزية لـ fudoki-hku

    fudoki-hku/FUDOKI

    76عرض على GitHub↗

    This repository is the official implementation of FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities.

    Discrete flow-based unified understanding and generation.

    Python
    عرض على GitHub↗76
  • jiyt17/rediffالصورة الرمزية لـ jiyt17

    jiyt17/ReDiff

    45عرض على GitHub↗

    We introduce ReDiff, a refining-enhanced vision-language diffusion model.

    Corrective framework for vision-language diffusion models.

    Python
    عرض على GitHub↗45
  • adobe-research/lavida-oالصورة الرمزية لـ adobe-research

    adobe-research/LaVida-O

    21عرض على GitHub↗

    [Paper](https://arxiv.org/abs/2509.19244) [Project Site](https://homepage.jackli.org/projects/lavida_o/index.html) [Huggingface](https://huggingface.co/jacklishufan/LaViDa-O-v1.0/tree/main)

    Elastic large masked diffusion for multimodal understanding and generation.

    Python
    عرض على GitHub↗21
  • zihohe/vidladaالصورة الرمزية لـ ziHoHe

    ziHoHe/VidLaDA

    10عرض على GitHub↗

    Bidirectional diffusion model for efficient video understanding.

    Python
    عرض على GitHub↗10
  1. Home
  2. Part of an Awesome List
  3. AI & Machine Learning
  4. Multimodal Diffusion Models