How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
TMLR23 Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks.
ECCV 2024 Best Paper Candidate & TPAMI 2025 PointLLM: Empowering Large Language Models to Understand Point Clouds
Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It functions as a design workspace where users can produce visual content and assets through a combination of local and cloud-based AI models. The project features a hybrid model orchestrator that routes requests between local model runners and remote APIs to balance data privacy with processing performance. It utilizes an infinite canvas collaborative tool for organizing storyboards and assets, and includes an image prompt optimizer to translate rough ideas into detailed generati
Emu Series: Generative Multimodal Models from BAAI
The main features of baaivision/emu are: Foundation Models, In Context Learning, Multimodal Agents.
Projects with overlapping indexed features include: openrobotlab/pointllm — [ECCV 2024 Best Paper Candidate & TPAMI 2025] PointLLM: Empowering Large Language Models to Understand Point Clouds. microsoft/mm-react — Official repo for MM-REACT. mshukor/unival — [TMLR23] Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks. 11cafe/jaaz — Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It… othersideai/self-operating-computer — This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs… qwenlm/qwen3-omni — Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video…