TMLR23 Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks.
Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. (Correspondence )
Emu Series: Generative Multimodal Models from BAAI
Monkey (LMM): Image Resolution and Text Label Are Important Things for Large Multi-modal Models (CVPR 2024 Highlight)
[ECCV 2024 Best Paper Candidate & TPAMI 2025] PointLLM: Empowering Large Language Models to Understand Point Clouds
openrobotlab/pointllm की मुख्य विशेषताएं हैं: 3D Scene Understanding, Foundation Models, Multimodal Agents, Multimodal Learning, Specialized Multimodal Tasks।
openrobotlab/pointllm के ओपन-सोर्स विकल्पों में शामिल हैं: next-gpt/next-gpt — Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. (Correspondence ). yuliang-liu/monkey — Monkey (LMM): Image Resolution and Text Label Are Important Things for Large Multi-modal Models (CVPR 2024 Highlight). baaivision/emu — Emu Series: Generative Multimodal Models from BAAI. mshukor/unival — [TMLR23] Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks. 11cafe/jaaz — Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It… othersideai/self-operating-computer — This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs…