awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
mshukor avatar

mshukor/UnIVAL

0
View on GitHub↗
236 estrellas·22 forks·Jupyter Notebook·Apache-2.0·6 vistas

UnIVAL

[TMLR23] Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks.

Features

  • Foundation Models - Unified architecture for image, video, audio, and language.
  • Multimodal Agents - Unified model for image, video, audio, and language tasks.

Historial de estrellas

Gráfico del historial de estrellas de mshukor/univalGráfico del historial de estrellas de mshukor/unival

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a UnIVAL

Proyectos open-source similares, clasificados según cuántas características comparten con UnIVAL.
  • openrobotlab/pointllmAvatar de OpenRobotLab

    OpenRobotLab/PointLLM

    1,026Ver en GitHub↗

    ECCV 2024 Best Paper Candidate & TPAMI 2025 PointLLM: Empowering Large Language Models to Understand Point Clouds

    Python
    Ver en GitHub↗1,026
  • baaivision/emuAvatar de baaivision

    baaivision/Emu

    1,775Ver en GitHub↗

    Emu Series: Generative Multimodal Models from BAAI

    Python
    Ver en GitHub↗1,775
  • 11cafe/jaazAvatar de 11cafe

    11cafe/jaaz

    6,384Ver en GitHub↗

    Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It functions as a design workspace where users can produce visual content and assets through a combination of local and cloud-based AI models. The project features a hybrid model orchestrator that routes requests between local model runners and remote APIs to balance data privacy with processing performance. It utilizes an infinite canvas collaborative tool for organizing storyboards and assets, and includes an image prompt optimizer to translate rough ideas into detailed generati

    TypeScript
    Ver en GitHub↗6,384
  • othersideai/self-operating-computerAvatar de OthersideAI

    OthersideAI/self-operating-computer

    10,153Ver en GitHub↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    Ver en GitHub↗10,153
Ver las 30 alternativas a UnIVAL→

Preguntas frecuentes

¿Qué hace mshukor/unival?

[TMLR23] Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks.

¿Cuáles son las características principales de mshukor/unival?

Las características principales de mshukor/unival son: Foundation Models, Multimodal Agents.

¿Qué alternativas de código abierto existen para mshukor/unival?

Las alternativas de código abierto para mshukor/unival incluyen: openrobotlab/pointllm — [ECCV 2024 Best Paper Candidate & TPAMI 2025] PointLLM: Empowering Large Language Models to Understand Point Clouds. baaivision/emu — Emu Series: Generative Multimodal Models from BAAI. qwenlm/qwen3-omni — Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video… othersideai/self-operating-computer — This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs… 11cafe/jaaz — Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It… simular-ai/agent-s — Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through…