20 open-source projects similar to internrobotics/g2vlm, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.
`bibtex @inproceedings{linghu20263d, title={3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding}, author={Linghu, Xiongkun and Huang, Jiangyong and Jia, Baoxiong and Huang, Siyuan}, booktitle={International Conference on Machine Learning}, year={2026} } `
This is a repo for paper "Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes". paper, project page
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
An Embodied Generalist Agent in 3D World
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
[📖 arXiv](https://arxiv.org/abs/2412.01292) [🤖 model](https://huggingface.co/Hoyard/LSceneLLM) [📑 dataset](https://huggingface.co/datasets/Hoyard/XR-Scene)
Official repository for the paper "Exploring the Potential of Encoder-free Architectures in 3D LMMs".
Official implementation of 'ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding with GPT and Prototype Guidance'.
CVPR 2025 The code for paper ''Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding''.
ECCV 2024 Best Paper Candidate & TPAMI 2025 PointLLM: Empowering Large Language Models to Understand Point Clouds
CVPR'24 Highlight GPT4Point: A Unified Framework for Point-Language Understanding and Generation.
Open-Vocabulary 3D Localization. Locate anything with natural language dialog! - Interactive Grounding. Humans will be able to chat with an agent to localize novel objects.
More Text, Less Point: Towards 3D Data-Efficient Point-Language Understanding Yuan Tang  Xu Han  Xianzhi Li ✝   Qiao Yu  Jinfeng Xu  Yixue Hao  Long Hu  Min Chen Huazhong University of Science and Technology South China University of Technology
3D-LLM: Injecting the 3D World into Large Language Models (NeurIPS 2023 Spotlight) Yining Hong , Haoyu Zhen , Peihao Chen , Shuhong Zheng , Yilun Du , Zhenfang Chen , Chuang Gan
This is an official repo for paper "Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning", ICCV 2025. paper
We build a multi-modal large language model for 3D scene understanding, excelling in tasks such as 3D grounding, captioning, and question answering.