awesome-repositories.com
Blog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

7 dépôts

Awesome GitHub RepositoriesGUI Agents

Agents capable of operating desktop and mobile user interfaces.

Explore 7 awesome GitHub repositories matching part of an awesome list · GUI Agents. Refine with filters or upvote what's useful.

Awesome GUI Agents GitHub Repositories

Trouvez les meilleurs dépôts grâce à l'IA.Nous recherchons les dépôts les plus pertinents grâce à l'IA.
  • microsoft/omniparserAvatar de microsoft

    microsoft/OmniParser

    24,377Voir sur GitHub↗

    OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual screen information to execute complex, multi-step tasks across operating system environments by bridging visual interface perception with language models. Through a continuous cycle of observation and command execution, the system grounds high-level natural language instructions into precise, coordinate-based actions. The project distinguishes itself by utilizing vision-based parsing to interact with software interfaces without requiring access to underlying application progr

    Vision-based parsing for GUI-driven agent automation.

    Jupyter Notebook
    Voir sur GitHub↗24,377
  • bytedance/ui-tarsAvatar de bytedance

    bytedance/UI-TARS

    9,622Voir sur GitHub↗

    UI-TARS is an LLM GUI automation framework and multimodal action grounding system. It functions as a GUI agent orchestrator and cross-platform device controller that uses large language models to interpret graphical interfaces and execute actions across desktop and mobile operating systems. The system translates model-generated coordinates into precise screen positions to interact with visual user interface elements. It employs a multimodal approach to interpret screen layouts and decomposes complex goals into multi-step trajectories through reasoning and error correction. The project provid

    Operates as a multimodal agent capable of interpreting screen layouts and performing multi-step tasks across desktop and mobile interfaces.

    Pythonresearch
    Voir sur GitHub↗9,622
  • microsoft/ufoAvatar de microsoft

    microsoft/UFO

    9,017Voir sur GitHub↗

    UFO is a multi-device task orchestrator and LLM agent orchestration framework designed to decompose natural language requests into executable task graphs. It functions as a cross-platform UI automation tool capable of performing interactions on Windows and mobile devices while routing tasks to distributed agents based on their hardware and software capabilities. The system is distinguished by its RAG-enhanced agent architecture, which integrates external documentation and previous execution traces to improve decision-making. It employs a hybrid UI detection approach that combines computer vis

    UI-focused agent for Windows operating system interaction.

    Pythonagentautomationcopilot
    Voir sur GitHub↗9,017
  • x-plug/mobileagentAvatar de X-PLUG

    X-PLUG/MobileAgent

    7,218Voir sur GitHub↗

    MobileAgent is an LLM-powered mobile automation agent and framework designed to navigate mobile user interfaces and execute multi-step tasks. It functions as a device interface automation system that maps semantic commands to screen coordinates to perform input events across mobile operating systems. The project operates as a cross-app workflow orchestrator, switching between native on-screen interface actions and external API tools to complete sophisticated operations. It includes a visual grounding system that analyzes screenshots and interface metadata to identify elements and validate the

    Autonomous multimodal agent for mobile device interaction.

    Pythonagentandroidapp
    Voir sur GitHub↗7,218
  • tencentqqgylab/appagentT

    TencentQQGYLab/AppAgent

    6,786Voir sur GitHub↗

    AppAgent is an autonomous system and Android app controller that uses large language models to navigate and execute tasks within mobile applications. It functions as a mobile UI automator and element mapper, capable of performing specific application tasks by utilizing documented user interface patterns and screen navigation. The framework differentiates itself through its ability to map application navigation and generate UI documentation via autonomous exploration or human-in-the-loop demonstrations. It employs a visual-language model to process screen screenshots and UI hierarchies to dete

    Multimodal agent framework for smartphone application usage.

    Python
    Voir sur GitHub↗6,786
  • thudm/cogvlmAvatar de THUDM

    THUDM/CogVLM

    6,742Voir sur GitHub↗

    CogVLM is a multimodal large language model designed to integrate visual and textual data for reasoning about images and generating natural language. It functions as a visual question answering system that analyzes image content to provide detailed descriptions or answer specific questions. The project includes a visual grounding model capable of mapping text descriptions to precise bounding box coordinates within an image. It also features a vision-based automation agent that analyzes screen captures to generate execution plans and interaction coordinates for software interfaces. The system

    Provides a vision-based agent capable of analyzing screens and interacting with desktop and mobile interfaces.

    Python
    Voir sur GitHub↗6,742
  • zai-org/cogvlmAvatar de zai-org

    zai-org/CogVLM

    6,742Voir sur GitHub↗

    CogVLM is a multimodal large language model designed for visual reasoning and multi-turn dialogue. It functions as a visual grounding model and a quantized vision model, combining text and image processing to perform complex understanding and maintain context across visual inputs. The project includes capabilities as a GUI automation agent, allowing it to analyze application screenshots, plan operational steps, and return precise screen coordinates for interface interaction. It further supports visual grounding by generating bounding box coordinates to map text descriptions to specific spatia

    Functions as an agent capable of operating and automating desktop and mobile user interfaces.

    Pythoncross-modalitylanguage-modelmulti-modal
    Voir sur GitHub↗6,742
  1. Home
  2. Part of an Awesome List
  3. AI & Machine Learning
  4. GUI Agents