awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Gen-Verse avatar

Gen-Verse/MMaDA

0
View on GitHub↗
1,656 stars·89 forks·Python·MIT·7 viewsopenreview.net/forum?id=wczmXLuLGd↗

MMaDA

Multimodal Large Diffusion Language Models (NeurIPS 2025)

Features

  • Multimodal Diffusion Models - Multimodal large diffusion language model architecture.

Star history

Star history chart for gen-verse/mmadaStar history chart for gen-verse/mmada

How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to MMaDA

Similar open-source projects, ranked by how many features they share with MMaDA.
  • ml-gsai/lladaML-GSAI avatar

    ML-GSAI/LLaDA

    3,580View on GitHub↗

    LLaDA is a masked diffusion language model and conditional text generator. It generates text by iteratively refining masked tokens through a diffusion process rather than predicting the next token in a sequence. The project functions as a vision-language diffusion model, converting visual inputs into text responses. It also serves as a preference optimization framework that uses log-likelihood estimation and evidence lower bounds to tune model responses. The system supports multi-round conversational AI and text sequence evaluation. It integrates vision-language embedding for cross-modal con

    Python
    View on GitHub↗3,580
  • vectorspacelab/omnigenVectorSpaceLab avatar

    VectorSpaceLab/OmniGen

    4,326View on GitHub↗

    OmniGen is a unified image generation model and diffusion framework that processes text, images, and vision tasks through a single system. It functions as a multimodal diffusion framework that treats diverse vision operations as unified image synthesis problems using shared model weights, removing the need for external adapter modules. The system supports subject-driven image generation to preserve the identity of objects from reference photos and allows for multi-reference image synthesis. It also operates as an instruction-based image editor, modifying visual content through natural languag

    Jupyter Notebookdiffusionimageimage-edit
    View on GitHub↗4,326
  • alpha-vllm/lumina-dimooAlpha-VLLM avatar

    Alpha-VLLM/Lumina-DiMOO

    1,001View on GitHub↗

    Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

    Python
    View on GitHub↗1,001
  • fudoki-hku/fudokifudoki-hku avatar

    fudoki-hku/FUDOKI

    76View on GitHub↗

    This repository is the official implementation of FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities.

    Python
    View on GitHub↗76
See all 15 alternatives to MMaDA→

Frequently asked questions

What does gen-verse/mmada do?

Multimodal Large Diffusion Language Models (NeurIPS 2025)

What are the main features of gen-verse/mmada?

The main features of gen-verse/mmada are: Multimodal Diffusion Models.

What are some open-source alternatives to gen-verse/mmada?

Open-source alternatives to gen-verse/mmada include: ml-gsai/llada — LLaDA is a masked diffusion language model and conditional text generator. It generates text by iteratively refining… vectorspacelab/omnigen — OmniGen is a unified image generation model and diffusion framework that processes text, images, and vision tasks… alpha-vllm/lumina-dimoo — Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding. hustvl/diffusionvl — DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models. jacklishufan/lavida — [[Paper]](paper/paper.pdf) [[Arxiv]](https://arxiv.org/abs/2505.16839) [[Checkpoints]](https://huggingface.co/collectio… fudoki-hku/fudoki — This repository is the official implementation of FUDOKI: Discrete Flow-based Unified Understanding and Generation via…