awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Meituan-AutoML avatar

Meituan-AutoML/MobileVLM

0
View on GitHub↗
1,359 estrellas·88 forks·Python·Apache-2.0·4 vistas

MobileVLM

MobileVLM: Vision Language Model for Mobile Devices

Features

  • Multimodal Agents - Fast vision-language assistant for mobile devices.
  • Vision Language Model - Listed in the “Vision Language Model” section of the Ailia Models awesome list.

Historial de estrellas

Gráfico del historial de estrellas de meituan-automl/mobilevlmGráfico del historial de estrellas de meituan-automl/mobilevlm

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a MobileVLM

Proyectos open-source similares, clasificados según cuántas características comparten con MobileVLM.
  • othersideai/self-operating-computerAvatar de OthersideAI

    OthersideAI/self-operating-computer

    10,153Ver en GitHub↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    Ver en GitHub↗10,153
  • qwenlm/qwen3-omniAvatar de QwenLM

    QwenLM/Qwen3-Omni

    3,843Ver en GitHub↗

    Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video within a single unified neural architecture. It functions as a real-time voice assistant and multimodal AI agent capable of reasoning across different media types and executing external tool-calling functions via APIs. The system supports low-latency conversational AI through autoregressive token streaming and natural turn-taking. It enables multilingual speech translation and generation across dozens of languages, featuring customizable speaker profiles and tones. The model's cap

    Jupyter Notebook
    Ver en GitHub↗3,843
  • 11cafe/jaazAvatar de 11cafe

    11cafe/jaaz

    6,384Ver en GitHub↗

    Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It functions as a design workspace where users can produce visual content and assets through a combination of local and cloud-based AI models. The project features a hybrid model orchestrator that routes requests between local model runners and remote APIs to balance data privacy with processing performance. It utilizes an infinite canvas collaborative tool for organizing storyboards and assets, and includes an image prompt optimizer to translate rough ideas into detailed generati

    TypeScript
    Ver en GitHub↗6,384
  • simular-ai/agent-sAvatar de simular-ai

    simular-ai/Agent-S

    11,855Ver en GitHub↗

    Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov

    Pythonagent-computer-interfaceai-agentscomputer-automation
    Ver en GitHub↗11,855
Ver las 30 alternativas a MobileVLM→

Preguntas frecuentes

¿Qué hace meituan-automl/mobilevlm?

MobileVLM: Vision Language Model for Mobile Devices

¿Cuáles son las características principales de meituan-automl/mobilevlm?

Las características principales de meituan-automl/mobilevlm son: Multimodal Agents, Vision Language Model.

¿Qué alternativas de código abierto existen para meituan-automl/mobilevlm?

Las alternativas de código abierto para meituan-automl/mobilevlm incluyen: qwenlm/qwen3-omni — Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video… simular-ai/agent-s — Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through… 11cafe/jaaz — Jaaz is a self-hosted AI design suite and multimodal workspace used for generating and editing images and videos. It… othersideai/self-operating-computer — This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs… cliport/cliport — CLIPort: What and Where Pathways for Robotic Manipulation Mohit Shridhar, Lucas Manuelli, Dieter Fox CoRL 2021. ai-chef/hugginggpt — The mission of JARVIS is to explore artificial general intelligence (AGI) and deliver cutting-edge research to the…