30 open-source projects similar to tulerfeng/video-r1, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.
The official repo for "Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models".
x 2025/09/26:🔥🔥🔥 We release our VideoChat-R1.5 model at Huggingface, paper, and eval code. - x 2025/09/22: 🎉🎉🎉 Our VideoChat-R1.5 is accepted by NIPS2025. - x 2025/04/22:🔥🔥🔥 We release our VideoChat-R1-caption at Huggingface. - x 2025/04/14:🔥🔥🔥 We release our VideoChat-R1 and…
VLM-R1 is a reasoning vision-language model and embodied AI framework designed to map visual inputs and language instructions into physical navigation waypoints and robotic actions. It functions as a multimodal policy optimizer and an open vocabulary detector capable of locating objects based on arbitrary natural language descriptions. The system distinguishes itself through the use of chain-of-thought reasoning and reinforcement learning to solve complex visual and spatial tasks. It utilizes a video semantic memory system, which employs a visual cache to maintain a history of live video for
Visual-RFT: Visual Reinforcement Fine-Tuning Ziyu Liu · Zeyi Sun · Yuhang Zang · Xiaoyi Dong · Yuhang Cao · Haodong Duan · Dahua Lin · Jiaqi Wang Accepted By ICCV 2025! 📖 Paper | 🤗 Datasets | 🤗 Daily Paper 🌈We introduce Visual Reinforcement Fine-tuning (Visual-RFT) , the first comprehensive…
Qwen2.5 is a suite of large language model foundation models designed for natural language generation, code production, and complex mathematical reasoning. The project encompasses a multilingual language model capable of processing dozens of languages and a specialized code generation model for technical problem solving and debugging. The framework is distinguished by its long context capabilities, enabling the analysis of massive inputs ranging from 256K up to 1 million tokens. It further functions as an agentic framework, utilizing standardized templates and parsers to execute autonomous wo
Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec
Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency
Exploring Applications of GRPO
An Open Large Reasoning Model for Real-World Solutions
Yi-Hsin Hung 1 , Yueqi Duan 1 , Equal Contribution. 1 Tsinghua University NeurIPS 2025 (Spotlight)
Project Page For "Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement"
TPAMI 2026 Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
Grounded Reasoning wiht Texts and Images (GRIT) is a novel method for training Multimodal Large Language Models (MLLMs) to perform grounded reasoning by generating reasoning chains that interleave natural language and explicit bounding box coordinates. This approach can use as few as 20 training…
Checkpoints take up a lot of space. Please email yninghong@gmail.com if you need them.
🧐 About | 🚀 Quick Start | 🐣 Agentless Mini | 📝 Citation | 🙏 Acknowledgements
R1-onevision, a visual language model capable of deep CoT reasoning.
OpenSeek aims to unite the global open source community to drive collaborative innovation in algorithms, data and systems to develop next-generation models.
DeepSeek-R1 is an open-weights large language model focused on advanced reasoning. It uses chain-of-thought processing and internal monologues to solve complex mathematical and logical problems by breaking tasks into sequential, verifiable thought processes. The model is developed using reinforcement learning to optimize reasoning patterns and verify logical steps. It employs a distillation process to transfer these high-performance logic capabilities from a large teacher model into smaller, computationally efficient versions. The training framework incorporates group relative policy optimiz