How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
Official Repo For OMG-LLaVA and OMG-Seg codebase CVPR-24 and NeurIPS-24
Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.
ECCV 2024 Best Paper Candidate & TPAMI 2025 PointLLM: Empowering Large Language Models to Understand Point Clouds
The official implement of VITA, VITA15, LongVITA, VITA-Audio, VITA-VLA, and VITA-E.
Project Page for "LISA: Reasoning Segmentation via Large Language Model"
The main features of dvlab-research/lisa are: Specialized Multimodal Tasks.
Projects with overlapping indexed features include: lxtgh/omg-seg — Official Repo For OMG-LLaVA and OMG-Seg codebase [CVPR-24 and NeurIPS-24]. microsoft/llava-med — Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities. openrobotlab/pointllm — [ECCV 2024 Best Paper Candidate & TPAMI 2025] PointLLM: Empowering Large Language Models to Understand Point Clouds. tencent/vita — The official implement of VITA, VITA15, LongVITA, VITA-Audio, VITA-VLA, and VITA-E. wentaoyuan/robopoint — A Vision-Language Model for Spatial Affordance Prediction in Robotics. yuliang-liu/monkey — Monkey (LMM): Image Resolution and Text Label Are Important Things for Large Multi-modal Models (CVPR 2024 Highlight).