awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to alpha-vllm/lumina-dimoo

Open-source alternatives to Lumina DiMOO

29 open-source projects similar to alpha-vllm/lumina-dimoo, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Lumina DiMOO alternative.

  • meituan-longcat/longcat-nextAvatar von meituan-longcat

    meituan-longcat/LongCat-Next

    442Auf GitHub ansehen↗
    Auf GitHub ansehen↗442
  • deepseek-ai/janusAvatar von deepseek-ai

    deepseek-ai/Janus

    17,746Auf GitHub ansehen↗

    Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec

    Pythonany-to-anyfoundation-modelsllm
    Auf GitHub ansehen↗17,746
  • mm-mvr/starM

    mm-mvr/star

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • opengvlab/internvl-uAvatar von OpenGVLab

    OpenGVLab/InternVL-U

    291Auf GitHub ansehen↗

    InternVL-U is a 4B-parameter unified multimodal model (UMM) that brings multimodal understanding, reasoning, image generation, image editing into a single framework.

    Python
    Auf GitHub ansehen↗291
  • facebookresearch/tuna-2Avatar von facebookresearch

    facebookresearch/tuna-2

    725Auf GitHub ansehen↗

    Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation

    Python
    Auf GitHub ansehen↗725

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Find more with AI search
  • oppomklab/u-llavaO

    OPPOMKLab/u-LLaVA

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • pku-yuangroup/uaeP

    PKU-YuanGroup/UAE

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • rainbowluocs/openomniR

    RainBowLuoCS/OpenOmni

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • skyworkai/unipicS

    SkyworkAI/UniPic

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • skyworkai/vitronAvatar von SkyworkAI

    SkyworkAI/Vitron

    577Auf GitHub ansehen↗

    NeurIPS 2024 Paper

    Python
    Auf GitHub ansehen↗577
  • bytedance/lanceAvatar von bytedance

    bytedance/Lance

    1,250Auf GitHub ansehen↗

    A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.

    Python
    Auf GitHub ansehen↗1,250
  • lehduong/onediffusionL

    lehduong/OneDiffusion

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • byteflow-ai/tokenflowAvatar von ByteFlow-AI

    ByteFlow-AI/TokenFlow

    465Auf GitHub ansehen↗

    CVPR 2025 🔥 Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation".

    Python
    Auf GitHub ansehen↗465
  • ml-gsai/lladaAvatar von ML-GSAI

    ML-GSAI/LLaDA

    3,580Auf GitHub ansehen↗

    LLaDA is a masked diffusion language model and conditional text generator. It generates text by iteratively refining masked tokens through a diffusion process rather than predicting the next token in a sequence. The project functions as a vision-language diffusion model, converting visual inputs into text responses. It also serves as a preference optimization framework that uses log-likelihood estimation and evidence lower bounds to tune model responses. The system supports multi-round conversational AI and text sequence evaluation. It integrates vision-language embedding for cross-modal con

    Python
    Auf GitHub ansehen↗3,580
  • vectorspacelab/omnigenAvatar von VectorSpaceLab

    VectorSpaceLab/OmniGen

    4,326Auf GitHub ansehen↗

    OmniGen is a unified image generation model and diffusion framework that processes text, images, and vision tasks through a single system. It functions as a multimodal diffusion framework that treats diverse vision operations as unified image synthesis problems using shared model weights, removing the need for external adapter modules. The system supports subject-driven image generation to preserve the identity of objects from reference photos and allows for multi-reference image synthesis. It also operates as an instruction-based image editor, modifying visual content through natural languag

    Jupyter Notebookdiffusionimageimage-edit
    Auf GitHub ansehen↗4,326
  • zihohe/vidladaAvatar von ziHoHe

    ziHoHe/VidLaDA

    10Auf GitHub ansehen↗
    Python
    Auf GitHub ansehen↗10
  • alexanderswerdlow/unidiscAvatar von alexanderswerdlow

    alexanderswerdlow/unidisc

    141Auf GitHub ansehen↗

    Unified Multimodal Discrete Diffusion

    Python
    Auf GitHub ansehen↗141
  • bytedance-seed/veomniB

    ByteDance-Seed/VeOmni

    0Auf GitHub ansehen↗
    Auf GitHub ansehen↗0
  • fudoki-hku/fudokiAvatar von fudoki-hku

    fudoki-hku/FUDOKI

    76Auf GitHub ansehen↗

    This repository is the official implementation of FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities.

    Python
    Auf GitHub ansehen↗76
  • gen-verse/mmadaAvatar von Gen-Verse

    Gen-Verse/MMaDA

    1,656Auf GitHub ansehen↗

    Multimodal Large Diffusion Language Models (NeurIPS 2025)

    Python
    Auf GitHub ansehen↗1,656
  • hustvl/diffusionvlAvatar von hustvl

    hustvl/DiffusionVL

    149Auf GitHub ansehen↗

    DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models

    Python
    Auf GitHub ansehen↗149
  • jacklishufan/lavidaAvatar von jacklishufan

    jacklishufan/LaViDa

    221Auf GitHub ansehen↗

    [Paper](paper/paper.pdf) [Arxiv](https://arxiv.org/abs/2505.16839) [Checkpoints](https://huggingface.co/collections/jacklishufan/lavida-10-682ecf5a5fa8c5df85c61ded) [Data](https://huggingface.co/datasets/jacklishufan/lavida-train) [Website](https://homepage.jackli.org/projects/lavida/)

    Python
    Auf GitHub ansehen↗221
  • jiyt17/rediffAvatar von jiyt17

    jiyt17/ReDiff

    45Auf GitHub ansehen↗

    We introduce ReDiff, a refining-enhanced vision-language diffusion model.

    Python
    Auf GitHub ansehen↗45
  • m-e-agi-lab/mudditAvatar von M-E-AGI-Lab

    M-E-AGI-Lab/Muddit

    117Auf GitHub ansehen↗

    Muddit is the 2nd generation Meissonic. It is built upon discrete diffusion for unified and efficient multimodal generation.

    Python
    Auf GitHub ansehen↗117
  • ml-gsai/llada-vAvatar von ML-GSAI

    ML-GSAI/LLaDA-V

    345Auf GitHub ansehen↗

    2026.03.23 We are excited to introduce LLaDA-o, the latest model in the LLaDA series. As an effective and length-adaptive omni diffusion model for unified multimodal understanding and generation, LLaDA-o extends the LLaDA line to broader multimodal settings, supporting visual understanding,…

    Python
    Auf GitHub ansehen↗345
  • openhelix-team/unified-diffusion-vlaAvatar von OpenHelix-Team

    OpenHelix-Team/Unified-Diffusion-VLA

    182Auf GitHub ansehen↗

    Jiayi Chen¹\,Wenxuan Song¹†\, Pengxiang Ding²˒³, Ziyang Zhou¹, Han Zhao²˒³, Feilong Tang⁴,Donglin Wang², Haoang Li¹‡

    Python
    Auf GitHub ansehen↗182
  • tyfeld/mmada-parallelAvatar von tyfeld

    tyfeld/MMaDA-Parallel

    300Auf GitHub ansehen↗

    ICLR 2026 MMaDA-Parallel: Parallel Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation

    Python
    Auf GitHub ansehen↗300
  • yu-rp/dimpleAvatar von yu-rp

    yu-rp/Dimple

    117Auf GitHub ansehen↗

    Dimple, the first Discrete Diffusion Multimodal Large Language Model

    Python
    Auf GitHub ansehen↗117
  • adobe-research/lavida-oAvatar von adobe-research

    adobe-research/LaVida-O

    21Auf GitHub ansehen↗

    [Paper](https://arxiv.org/abs/2509.19244) [Project Site](https://homepage.jackli.org/projects/lavida_o/index.html) [Huggingface](https://huggingface.co/jacklishufan/LaViDa-O-v1.0/tree/main)

    Python
    Auf GitHub ansehen↗21