awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to eric-ai-lab/grit

Projects sharing features with GRIT

30 open-source projects similar to eric-ai-lab/grit, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.

  • deepseek-ai/janusdeepseek-ai avatar

    deepseek-ai/Janus

    17,746View on GitHub↗

    Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec

    Pythonany-to-anyfoundation-modelsllm
    View on GitHub↗17,746
  • diankun-wu/spatial-mllmdiankun-wu avatar

    diankun-wu/Spatial-MLLM

    470View on GitHub↗

    Yi-Hsin Hung 1 , Yueqi Duan 1 , Equal Contribution. 1 Tsinghua University NeurIPS 2025 (Spotlight)

    Python
    View on GitHub↗470
  • egolife-ai/ego-r1egolife-ai avatar

    egolife-ai/Ego-R1

    158View on GitHub↗

    TPAMI 2026 Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning

    Python
    View on GitHub↗158
  • hitsz-tmg/veripoHITsz-TMG avatar

    HITsz-TMG/VerIPO

    10View on GitHub↗

    📄 Paper Link 🤗 VerIPO-7B-v1.0

    Python
    View on GitHub↗10
  • jun297/v1jun297 avatar

    jun297/v1

    20View on GitHub↗

    Jiwan Chung   Junhyeok Kim   Siyeol Kim   Jaeyoung Lee   Minsoo Kim   Youngjae Yu

    Python
    View on GitHub↗20

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • kwai-keye/keyeKwai-Keye avatar

    Kwai-Keye/Keye

    797View on GitHub↗
    Python
    View on GitHub↗797
  • liuziyu77/visual-rftLiuziyu77 avatar

    Liuziyu77/Visual-RFT

    2,250View on GitHub↗

    Visual-RFT: Visual Reinforcement Fine-Tuning Ziyu Liu · Zeyi Sun · Yuhang Zang · Xiaoyi Dong · Yuhang Cao · Haodong Duan · Dahua Lin · Jiaqi Wang Accepted By ICCV 2025! 📖 Paper | 🤗 Datasets | 🤗 Daily Paper 🌈We introduce Visual Reinforcement Fine-tuning (Visual-RFT) , the first comprehensive…

    Jupyter Notebook
    View on GitHub↗2,250
  • longmalongma/tw-grpolongmalongma avatar

    longmalongma/TW-GRPO

    36View on GitHub↗

    🤗 Model &nbsp&nbsp | &nbsp&nbsp 📑 Paper &nbsp&nbsp

    Python
    View on GitHub↗36
  • maifoundations/visionary-r1maifoundations avatar

    maifoundations/Visionary-R1

    44View on GitHub↗

    Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning A new RL method for visual reasoning, which significantly outperforms vanilla GRPO, and bypasses the need for explicit chain-of-thought supervision during training.

    Python
    View on GitHub↗44
  • nvlabs/long-rlN

    NVlabs/Long-RL

    0View on GitHub↗

    Scaling RL to Long Videos Paper Yukang Chen , Wei Huang , Baifeng Shi, Qinghao Hu, Hanrong Ye, Ligeng Zhu, Zhijian Liu, Pavlo Molchanov, Jan Kautz, Xiaojuan Qi, Sifei Liu,Hongxu Yin, Yao Lu, Song Han

    View on GitHub↗0
  • om-ai-lab/vlm-r1om-ai-lab avatar

    om-ai-lab/VLM-R1

    5,991View on GitHub↗

    VLM-R1 is a reasoning vision-language model and embodied AI framework designed to map visual inputs and language instructions into physical navigation waypoints and robotic actions. It functions as a multimodal policy optimizer and an open vocabulary detector capable of locating objects based on arbitrary natural language descriptions. The system distinguishes itself through the use of chain-of-thought reasoning and reinforcement learning to solve complex visual and spatial tasks. It utilizes a video semantic memory system, which employs a visual cache to maintain a history of live video for

    Python
    View on GitHub↗5,991
  • opengvlab/videochat-r1OpenGVLab avatar

    OpenGVLab/VideoChat-R1

    267View on GitHub↗

    x 2025/09/26:🔥🔥🔥 We release our VideoChat-R1.5 model at Huggingface, paper, and eval code. - x 2025/09/22: 🎉🎉🎉 Our VideoChat-R1.5 is accepted by NIPS2025. - x 2025/04/22:🔥🔥🔥 We release our VideoChat-R1-caption at Huggingface. - x 2025/04/14:🔥🔥🔥 We release our VideoChat-R1 and…

    Python
    View on GitHub↗267
  • osilly/vision-r1Osilly avatar

    Osilly/Vision-R1

    1,475View on GitHub↗

    The official repo for "Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models".

    Python
    View on GitHub↗1,475
  • ouyangkun10/spacerOuyangKun10 avatar

    OuyangKun10/SpaceR

    112View on GitHub↗

    📖 Paper 🤗 SpaceR 📊 SpaceR-151k

    Python
    View on GitHub↗112
  • pzyseere/metaspatialPzySeere avatar

    PzySeere/MetaSpatial

    209View on GitHub↗

    MetaSpatial enhances spatial reasoning in VLMs using RL, internalizing 3D spatial reasoning to enable real-time 3D scene generation without hard-coded optimizations. By incorporating physics-aware constraints and rendering image evaluations, our framework optimizes layout coherence, physical…

    Python
    View on GitHub↗209
  • qiwang98/videorftQiWang98 avatar

    QiWang98/VideoRFT

    65View on GitHub↗

    2025/09/19 Our paper has been accepted to NeurIPS 2025 🎉! - 2025/06/01 We released our 3B Models (🤗VideoRFT-SFT-3B and 🤗VideoRFT-3B) to huggingface. - 2025/05/25 We released our 7B Models (🤗VideoRFT-SFT-7B and 🤗VideoRFT-7B) to huggingface. - 2025/05/20 We released our Datasets…

    Python
    View on GitHub↗65
  • tiger-ai-lab/pixel-reasonerTIGER-AI-Lab avatar

    TIGER-AI-Lab/Pixel-Reasoner

    298View on GitHub↗

    Haozhe Wang † , Weiming Ren , Fangzhen Lin , Wenhu Chen ‡ Equal Contribution. † Project Lead. ‡ Correspondence.

    Python
    View on GitHub↗298
  • tsinghuac3i/adsqaTsinghuaC3I avatar

    TsinghuaC3I/AdsQA

    36View on GitHub↗

    ICCV 2025 AdsQA: Towards Advertisement Video Understanding Arxiv: https://arxiv.org/abs/2509.08621

    Python
    View on GitHub↗36
  • tulerfeng/video-r1tulerfeng avatar

    tulerfeng/Video-R1

    878View on GitHub↗

    📖 Paper 🤗 Video-R1-7B-model 🤗 Video-R1-train-data 🤖 Video-R1-7B-model 🤖 Video-R1-train-data

    Python
    View on GitHub↗878
  • visual-agent/deepeyesVisual-Agent avatar

    Visual-Agent/DeepEyes

    1,239View on GitHub↗

    DeepEyes: Incentivizing “Thinking with Images” via Reinforcement Learning

    Python
    View on GitHub↗1,239
  • wangqinsi1/vision-zerowangqinsi1 avatar

    wangqinsi1/Vision-Zero

    136View on GitHub↗

    A domain-agnostic framework enabling VLM self-improvement through competitive visual games

    Python
    View on GitHub↗136
  • xtong-zhang/chain-of-focusxtong-zhang avatar

    xtong-zhang/Chain-of-Focus

    69View on GitHub↗

    Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs

    HTML
    View on GitHub↗69
  • yihedeng9/openvlthinkeryihedeng9 avatar

    yihedeng9/OpenVLThinker

    152View on GitHub↗

    Yihe Deng , Nanyun Peng , Kai-Wei Chang

    Python
    View on GitHub↗152
  • yix8/visualplanningyix8 avatar

    yix8/VisualPlanning

    362View on GitHub↗

    Visual Planning: Let's Think Only with Images If you find this project interesting, please give us a star ⭐ on GitHub to support us. 🙏🙏

    Python
    View on GitHub↗362
  • zhangquanchen/sifthinkerzhangquanchen avatar

    zhangquanchen/SIFThinker

    22View on GitHub↗

    SIFThinker: Spatially-Aware Image Focus for Visual Reasoning

    View on GitHub↗22
  • zhangquanchen/visrlzhangquanchen avatar

    zhangquanchen/VisRL

    46View on GitHub↗

    Visual understanding is inherently intention-driven—humans selectively focus on different regions of a scene based on their goals. Recent advances in large multimodal models (LMMs) enable flexible expression of such intentions through natural language, allowing queries to guide visual reasoning…

    Python
    View on GitHub↗46
  • zhaochen0110/openthinkimgzhaochen0110 avatar

    zhaochen0110/OpenThinkIMG

    392View on GitHub↗

    Use Vision Tools, Think with Images

    Jupyter Notebook
    View on GitHub↗392
  • zhijie-group/r1-zero-vsizhijie-group avatar

    zhijie-group/R1-Zero-VSI

    42View on GitHub↗

    Improved Visual-Spatial Reasoning via R1-Zero-Like Training Zhenyi Liao , Qingsong Xie , Yanhao Zhang , Zijian Kong , Haonan Lu , Zhenyu Yang , Zhijie Deng Datasets | 🤗 Daily Paper -->

    View on GitHub↗42
  • zhoues/roboreferZhoues avatar

    Zhoues/RoboRefer

    261View on GitHub↗

    RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics

    Python
    View on GitHub↗261
  • zzzhhzzz/ground-r1zzzhhzzz avatar

    zzzhhzzz/Ground-R1

    43View on GitHub↗

    Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning If you like our project, please give us a star ⭐ on GitHub for the latest update.

    Python
    View on GitHub↗43