](https://huggingface.co/Dream-org/Dream-v0-Base-7B)
Current Diffusion Language Models (DLMs) have been studied at a smaller scale compared to their autoregressive (AR) counterparts and lack fair comparison on language modeling benchmarks. Additionally, training diffusion models from scratch at scale remains challenging. We propose adapting…
Training Optimal Large Diffusion Language Models Jinjie Ni†, Qian Liu, Chao Du, Longxu Dou, Hang Yan, Zili Wang, Tianyu Pang, Michael Qizhe Shieh
We introduce LLaDA 1.5, a competitive large diffusion language model, trained by variance-reduced preference optimization (VRPO).
The main features of ml-gsai/llada-1.5 are: Language Diffusion Models, Training and Alignment.
Open-source alternatives to ml-gsai/llada-1.5 include: jinjieni/megadlms — MegaDLMs. jinjieni/quokka — Training Optimal Large Diffusion Language Models Jinjie Ni†, Qian Liu, Chao Du, Longxu Dou, Hang Yan, Zili Wang,… hkunlp/diffullama — Current Diffusion Language Models (DLMs) have been studied at a smaller scale compared to their autoregressive (AR)… hkunlp/dream — ](https://huggingface.co/Dream-org/Dream-v0-Base-7B). autonomousvision/mdpo — [[Paper]](https://arxiv.org/pdf/2508.13148) [[Project]](https://cli212.github.io/MDPO/). amap-ml/ar-map — Are Autoregressive Large Language Models Implicit Teachers for Diffusion Large Language Models? A comprehensive…