30 open-source projects similar to eric-ai-lab/grit, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.
Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec
Yi-Hsin Hung 1 , Yueqi Duan 1 , Equal Contribution. 1 Tsinghua University NeurIPS 2025 (Spotlight)
TPAMI 2026 Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
Jiwan Chung   Junhyeok Kim   Siyeol Kim   Jaeyoung Lee   Minsoo Kim   Youngjae Yu
Visual-RFT: Visual Reinforcement Fine-Tuning Ziyu Liu · Zeyi Sun · Yuhang Zang · Xiaoyi Dong · Yuhang Cao · Haodong Duan · Dahua Lin · Jiaqi Wang Accepted By ICCV 2025! 📖 Paper | 🤗 Datasets | 🤗 Daily Paper 🌈We introduce Visual Reinforcement Fine-tuning (Visual-RFT) , the first comprehensive…
Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning A new RL method for visual reasoning, which significantly outperforms vanilla GRPO, and bypasses the need for explicit chain-of-thought supervision during training.
Scaling RL to Long Videos Paper Yukang Chen , Wei Huang , Baifeng Shi, Qinghao Hu, Hanrong Ye, Ligeng Zhu, Zhijian Liu, Pavlo Molchanov, Jan Kautz, Xiaojuan Qi, Sifei Liu,Hongxu Yin, Yao Lu, Song Han
VLM-R1 is a reasoning vision-language model and embodied AI framework designed to map visual inputs and language instructions into physical navigation waypoints and robotic actions. It functions as a multimodal policy optimizer and an open vocabulary detector capable of locating objects based on arbitrary natural language descriptions. The system distinguishes itself through the use of chain-of-thought reasoning and reinforcement learning to solve complex visual and spatial tasks. It utilizes a video semantic memory system, which employs a visual cache to maintain a history of live video for
x 2025/09/26:🔥🔥🔥 We release our VideoChat-R1.5 model at Huggingface, paper, and eval code. - x 2025/09/22: 🎉🎉🎉 Our VideoChat-R1.5 is accepted by NIPS2025. - x 2025/04/22:🔥🔥🔥 We release our VideoChat-R1-caption at Huggingface. - x 2025/04/14:🔥🔥🔥 We release our VideoChat-R1 and…
The official repo for "Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models".
MetaSpatial enhances spatial reasoning in VLMs using RL, internalizing 3D spatial reasoning to enable real-time 3D scene generation without hard-coded optimizations. By incorporating physics-aware constraints and rendering image evaluations, our framework optimizes layout coherence, physical…
2025/09/19 Our paper has been accepted to NeurIPS 2025 🎉! - 2025/06/01 We released our 3B Models (🤗VideoRFT-SFT-3B and 🤗VideoRFT-3B) to huggingface. - 2025/05/25 We released our 7B Models (🤗VideoRFT-SFT-7B and 🤗VideoRFT-7B) to huggingface. - 2025/05/20 We released our Datasets…
Haozhe Wang † , Weiming Ren , Fangzhen Lin , Wenhu Chen ‡ Equal Contribution. † Project Lead. ‡ Correspondence.
ICCV 2025 AdsQA: Towards Advertisement Video Understanding Arxiv: https://arxiv.org/abs/2509.08621
📖 Paper 🤗 Video-R1-7B-model 🤗 Video-R1-train-data 🤖 Video-R1-7B-model 🤖 Video-R1-train-data
DeepEyes: Incentivizing “Thinking with Images” via Reinforcement Learning
A domain-agnostic framework enabling VLM self-improvement through competitive visual games
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
Visual Planning: Let's Think Only with Images If you find this project interesting, please give us a star ⭐ on GitHub to support us. 🙏🙏
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
Visual understanding is inherently intention-driven—humans selectively focus on different regions of a scene based on their goals. Recent advances in large multimodal models (LMMs) enable flexible expression of such intentions through natural language, allowing queries to guide visual reasoning…
Use Vision Tools, Think with Images
Improved Visual-Spatial Reasoning via R1-Zero-Like Training Zhenyi Liao , Qingsong Xie , Yanhao Zhang , Zijian Kong , Haonan Lu , Zhenyu Yang , Zhijie Deng Datasets | 🤗 Daily Paper -->
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning If you like our project, please give us a star ⭐ on GitHub for the latest update.