15 रिपॉजिटरी
Multimodal models capable of understanding and generating multiple media types.
Explore 15 awesome GitHub repositories matching part of an awesome list · Unified Models. Refine with filters or upvote what's useful.
Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec
Unified multimodal understanding and generation model.
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
Unified multimodal model for efficient processing.
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
Unified multimodal generation model.
Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation
Unified multimodal model for diverse applications.
NeurIPS 2024 Paper
Unified vision-language model for generation and editing.
CVPR 2025 🔥 Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation".
Unified framework for multimodal token processing.
InternVL-U is a 4B-parameter unified multimodal model (UMM) that brings multimodal understanding, reasoning, image generation, image editing into a single framework.
Unified vision-language model for diverse tasks.