29 open-source projects similar to alpha-vllm/lumina-dimoo, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Lumina DiMOO alternative.
Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec
InternVL-U is a 4B-parameter unified multimodal model (UMM) that brings multimodal understanding, reasoning, image generation, image editing into a single framework.
Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
CVPR 2025 🔥 Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation".
LLaDA is a masked diffusion language model and conditional text generator. It generates text by iteratively refining masked tokens through a diffusion process rather than predicting the next token in a sequence. The project functions as a vision-language diffusion model, converting visual inputs into text responses. It also serves as a preference optimization framework that uses log-likelihood estimation and evidence lower bounds to tune model responses. The system supports multi-round conversational AI and text sequence evaluation. It integrates vision-language embedding for cross-modal con
OmniGen is a unified image generation model and diffusion framework that processes text, images, and vision tasks through a single system. It functions as a multimodal diffusion framework that treats diverse vision operations as unified image synthesis problems using shared model weights, removing the need for external adapter modules. The system supports subject-driven image generation to preserve the identity of objects from reference photos and allows for multi-reference image synthesis. It also operates as an instruction-based image editor, modifying visual content through natural languag
This repository is the official implementation of FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities.
Multimodal Large Diffusion Language Models (NeurIPS 2025)
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
[Paper](paper/paper.pdf) [Arxiv](https://arxiv.org/abs/2505.16839) [Checkpoints](https://huggingface.co/collections/jacklishufan/lavida-10-682ecf5a5fa8c5df85c61ded) [Data](https://huggingface.co/datasets/jacklishufan/lavida-train) [Website](https://homepage.jackli.org/projects/lavida/)
We introduce ReDiff, a refining-enhanced vision-language diffusion model.
Muddit is the 2nd generation Meissonic. It is built upon discrete diffusion for unified and efficient multimodal generation.
2026.03.23 We are excited to introduce LLaDA-o, the latest model in the LLaDA series. As an effective and length-adaptive omni diffusion model for unified multimodal understanding and generation, LLaDA-o extends the LLaDA line to broader multimodal settings, supporting visual understanding,…
Jiayi Chen¹\,Wenxuan Song¹†\, Pengxiang Ding²˒³, Ziyang Zhou¹, Han Zhao²˒³, Feilong Tang⁴,Donglin Wang², Haoang Li¹‡
ICLR 2026 MMaDA-Parallel: Parallel Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
[Paper](https://arxiv.org/abs/2509.19244) [Project Site](https://homepage.jackli.org/projects/lavida_o/index.html) [Huggingface](https://huggingface.co/jacklishufan/LaViDa-O-v1.0/tree/main)