TMLR23 Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks.
Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. (Correspondence )
Emu Series: Generative Multimodal Models from BAAI
Monkey (LMM): Image Resolution and Text Label Are Important Things for Large Multi-modal Models (CVPR 2024 Highlight)
[ECCV 2024 Best Paper Candidate & TPAMI 2025] PointLLM: Empowering Large Language Models to Understand Point Clouds
Principalele funcționalități ale openrobotlab/pointllm sunt: 3D Scene Understanding, Foundation Models, Multimodal Agents, Multimodal Learning, Specialized Multimodal Tasks.
Alternativele open-source pentru openrobotlab/pointllm includ: next-gpt/next-gpt — Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. (Correspondence ). yuliang-liu/monkey — Monkey (LMM): Image Resolution and Text Label Are Important Things for Large Multi-modal Models (CVPR 2024 Highlight). baaivision/emu — Emu Series: Generative Multimodal Models from BAAI. mshukor/unival — [TMLR23] Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks. 11cafe/jaaz — Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It… othersideai/self-operating-computer — This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs…