awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
OpenHelix-Team avatar

OpenHelix-Team/Unified-Diffusion-VLA

0
View on GitHub↗
182 estrellas·9 forks·Python·MIT·4 vistas

Unified Diffusion VLA

Jiayi Chen¹,Wenxuan Song¹†, Pengxiang Ding²˒³, Ziyang Zhou¹, Han Zhao²˒³, Feilong Tang⁴,Donglin Wang², Haoang Li¹‡

Features

  • Multimodal Diffusion Models - Joint discrete denoising for vision-language-action models.

Historial de estrellas

Gráfico del historial de estrellas de openhelix-team/unified-diffusion-vlaGráfico del historial de estrellas de openhelix-team/unified-diffusion-vla

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a Unified Diffusion VLA

Proyectos open-source similares, clasificados según cuántas características comparten con Unified Diffusion VLA.
  • ml-gsai/lladaAvatar de ML-GSAI

    ML-GSAI/LLaDA

    3,580Ver en GitHub↗

    LLaDA is a masked diffusion language model and conditional text generator. It generates text by iteratively refining masked tokens through a diffusion process rather than predicting the next token in a sequence. The project functions as a vision-language diffusion model, converting visual inputs into text responses. It also serves as a preference optimization framework that uses log-likelihood estimation and evidence lower bounds to tune model responses. The system supports multi-round conversational AI and text sequence evaluation. It integrates vision-language embedding for cross-modal con

    Python
    Ver en GitHub↗3,580
  • vectorspacelab/omnigenAvatar de VectorSpaceLab

    VectorSpaceLab/OmniGen

    4,326Ver en GitHub↗

    OmniGen is a unified image generation model and diffusion framework that processes text, images, and vision tasks through a single system. It functions as a multimodal diffusion framework that treats diverse vision operations as unified image synthesis problems using shared model weights, removing the need for external adapter modules. The system supports subject-driven image generation to preserve the identity of objects from reference photos and allows for multi-reference image synthesis. It also operates as an instruction-based image editor, modifying visual content through natural languag

    Jupyter Notebookdiffusionimageimage-edit
    Ver en GitHub↗4,326
  • alpha-vllm/lumina-dimooAvatar de Alpha-VLLM

    Alpha-VLLM/Lumina-DiMOO

    1,001Ver en GitHub↗

    Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

    Python
    Ver en GitHub↗1,001
  • fudoki-hku/fudokiAvatar de fudoki-hku

    fudoki-hku/FUDOKI

    76Ver en GitHub↗

    This repository is the official implementation of FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities.

    Python
    Ver en GitHub↗76
Ver las 15 alternativas a Unified Diffusion VLA→

Preguntas frecuentes

¿Qué hace openhelix-team/unified-diffusion-vla?

Jiayi Chen¹\,Wenxuan Song¹†\, Pengxiang Ding²˒³, Ziyang Zhou¹, Han Zhao²˒³, Feilong Tang⁴,Donglin Wang², Haoang Li¹‡

¿Cuáles son las características principales de openhelix-team/unified-diffusion-vla?

Las características principales de openhelix-team/unified-diffusion-vla son: Multimodal Diffusion Models.

¿Qué alternativas de código abierto existen para openhelix-team/unified-diffusion-vla?

Las alternativas de código abierto para openhelix-team/unified-diffusion-vla incluyen: ml-gsai/llada — LLaDA is a masked diffusion language model and conditional text generator. It generates text by iteratively refining… vectorspacelab/omnigen — OmniGen is a unified image generation model and diffusion framework that processes text, images, and vision tasks… alpha-vllm/lumina-dimoo — Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding. gen-verse/mmada — Multimodal Large Diffusion Language Models (NeurIPS 2025). hustvl/diffusionvl — DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models. fudoki-hku/fudoki — This repository is the official implementation of FUDOKI: Discrete Flow-based Unified Understanding and Generation via…