16 repository-uri
Models integrating vision, language, and action through diffusion processes.
Explore 16 awesome GitHub repositories matching part of an awesome list · Multimodal Diffusion Models. Refine with filters or upvote what's useful.
OmniGen este un model unificat de generare a imaginilor și un framework de difuzie care procesează text, imagini și sarcini de viziune printr-un singur sistem. Funcționează ca un framework de difuzie multimodal care tratează diverse operațiuni de viziune ca probleme unificate de sinteză a imaginilor folosind ponderi de model partajate, eliminând nevoia de module adaptoare externe. Sistemul suportă generarea de imagini bazată pe subiect pentru a păstra identitatea obiectelor din fotografiile de referință și permite sinteza imaginilor multi-referință. De asemenea, operează ca un editor de imagini bazat pe instrucțiuni, modificând conținutul vizual prin prompt-uri în limbaj natural. Framework-ul se extinde la sarcini de viziune computațională generativă, unde operațiuni precum detectarea marginilor și recunoașterea posturii sunt efectuate prin transformarea lor în sarcini de sinteză. Performanța pe sarcini specifice poate fi îmbunătățită prin fine-tuning-ul ponderilor modelului și adaptare low-rank.
Provides a multimodal diffusion framework integrating vision and language via shared model weights.
LLaDA is a masked diffusion language model and conditional text generator. It generates text by iteratively refining masked tokens through a diffusion process rather than predicting the next token in a sequence. The project functions as a vision-language diffusion model, converting visual inputs into text responses. It also serves as a preference optimization framework that uses log-likelihood estimation and evidence lower bounds to tune model responses. The system supports multi-round conversational AI and text sequence evaluation. It integrates vision-language embedding for cross-modal con
A system that converts visual inputs into text responses using a diffusion process for complex multimodal tasks.
Multimodal Large Diffusion Language Models (NeurIPS 2025)
Multimodal large diffusion language model architecture.
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
Omni-diffusion model for multimodal generation and understanding.
2026.03.23 We are excited to introduce LLaDA-o, the latest model in the LLaDA series. As an effective and length-adaptive omni diffusion model for unified multimodal understanding and generation, LLaDA-o extends the LLaDA line to broader multimodal settings, supporting visual understanding,…
Large language diffusion models with visual instruction tuning.
ICLR 2026 MMaDA-Parallel: Parallel Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
Multimodal diffusion model for thinking-aware editing and generation.
[Paper](paper/paper.pdf) [Arxiv](https://arxiv.org/abs/2505.16839) [Checkpoints](https://huggingface.co/collections/jacklishufan/lavida-10-682ecf5a5fa8c5df85c61ded) [Data](https://huggingface.co/datasets/jacklishufan/lavida-train) [Website](https://homepage.jackli.org/projects/lavida/)
Large diffusion language model for multimodal understanding.
Jiayi Chen¹\,Wenxuan Song¹†\, Pengxiang Ding²˒³, Ziyang Zhou¹, Han Zhao²˒³, Feilong Tang⁴,Donglin Wang², Haoang Li¹‡
Joint discrete denoising for vision-language-action models.
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
Translating autoregressive models into vision-language diffusion models.
Unified Multimodal Discrete Diffusion
Unified multimodal discrete diffusion framework.
Dimple, the first Discrete Diffusion Multimodal Large Language Model
Discrete diffusion multimodal model with parallel decoding.
Muddit is the 2nd generation Meissonic. It is built upon discrete diffusion for unified and efficient multimodal generation.
Unified discrete diffusion for multimodal generation beyond text-to-image.
This repository is the official implementation of FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities.
Discrete flow-based unified understanding and generation.
We introduce ReDiff, a refining-enhanced vision-language diffusion model.
Corrective framework for vision-language diffusion models.
[Paper](https://arxiv.org/abs/2509.19244) [Project Site](https://homepage.jackli.org/projects/lavida_o/index.html) [Huggingface](https://huggingface.co/jacklishufan/LaViDa-O-v1.0/tree/main)
Elastic large masked diffusion for multimodal understanding and generation.