How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
x 2025/09/26:🔥🔥🔥 We release our VideoChat-R1.5 model at Huggingface, paper, and eval code. - x 2025/09/22: 🎉🎉🎉 Our VideoChat-R1.5 is accepted by NIPS2025. - x 2025/04/22:🔥🔥🔥 We release our VideoChat-R1-caption at Huggingface. - x 2025/04/14:🔥🔥🔥 We release our VideoChat-R1 and…
VLM-R1 is a reasoning vision-language model and embodied AI framework designed to map visual inputs and language instructions into physical navigation waypoints and robotic actions. It functions as a multimodal policy optimizer and an open vocabulary detector capable of locating objects based on arbitrary natural language descriptions. The system distinguishes itself through the use of chain-of-thought reasoning and reinforcement learning to solve complex visual and spatial tasks. It utilizes a video semantic memory system, which employs a visual cache to maintain a history of live video for
Visual-RFT: Visual Reinforcement Fine-Tuning Ziyu Liu · Zeyi Sun · Yuhang Zang · Xiaoyi Dong · Yuhang Cao · Haodong Duan · Dahua Lin · Jiaqi Wang Accepted By ICCV 2025! 📖 Paper | 🤗 Datasets | 🤗 Daily Paper 🌈We introduce Visual Reinforcement Fine-tuning (Visual-RFT) , the first comprehensive…
The official repo for "Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models".
[📖 Paper] [🤗 Video-R1-7B-model] [🤗 Video-R1-train-data] [🤖 Video-R1-7B-model] [🤖 Video-R1-train-data]
The main features of tulerfeng/video-r1 are: Multimodal Understanding, Reasoning Models.
Projects with overlapping indexed features include: opengvlab/videochat-r1 — [x] 2025/09/26:🔥🔥🔥 We release our VideoChat-R1.5 model at Huggingface, paper, and eval code. - [x] 2025/09/22:… osilly/vision-r1 — The official repo for "Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models". om-ai-lab/vlm-r1 — VLM-R1 is a reasoning vision-language model and embodied AI framework designed to map visual inputs and language… liuziyu77/visual-rft — Visual-RFT: Visual Reinforcement Fine-Tuning Ziyu Liu · Zeyi Sun · Yuhang Zang · Xiaoyi Dong · Yuhang Cao · Haodong… deepseek-ai/janus — Janus is a multimodal large language model and unified framework that integrates visual understanding and image… qwenlm/qwen2.5 — Qwen2.5 is a suite of large language model foundation models designed for natural language generation, code…