How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
[CVPR 2025] ๐ฅ Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation".
The main features of byteflow-ai/tokenflow are: Computer Vision Research, Generative AI, Unified Models, Unified Multimodal Models.
Open-source alternatives to byteflow-ai/tokenflow include: bytedance/x-dyna โ [CVPR 2025 Highlight] X-Dyna: Expressive Dynamic Human Image Animation. facebookresearch/tuna-2 โ Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation. alpha-vllm/lumina-dimoo โ Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding. bytedance/lance โ A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing. deepseek-ai/janus โ Janus is a multimodal large language model and unified framework that integrates visual understanding and imageโฆ hustvl/lightningdit โ [CVPR 2025 Oral] Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models.
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
CVPR 2025 Highlight X-Dyna: Expressive Dynamic Human Image Animation
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec