awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

7 个仓库

Awesome GitHub RepositoriesGUI Agents

Agents capable of operating desktop and mobile user interfaces.

Explore 7 awesome GitHub repositories matching part of an awesome list · GUI Agents. Refine with filters or upvote what's useful.

Awesome GUI Agents GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • microsoft/omniparsermicrosoft 的头像

    microsoft/OmniParser

    24,377在 GitHub 上查看↗

    OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual screen information to execute complex, multi-step tasks across operating system environments by bridging visual interface perception with language models. Through a continuous cycle of observation and command execution, the system grounds high-level natural language instructions into precise, coordinate-based actions. The project distinguishes itself by utilizing vision-based parsing to interact with software interfaces without requiring access to underlying application progr

    Vision-based parsing for GUI-driven agent automation.

    Jupyter Notebook
    在 GitHub 上查看↗24,377
  • bytedance/ui-tarsbytedance 的头像

    bytedance/UI-TARS

    9,622在 GitHub 上查看↗

    UI-TARS is an LLM GUI automation framework and multimodal action grounding system. It functions as a GUI agent orchestrator and cross-platform device controller that uses large language models to interpret graphical interfaces and execute actions across desktop and mobile operating systems. The system translates model-generated coordinates into precise screen positions to interact with visual user interface elements. It employs a multimodal approach to interpret screen layouts and decomposes complex goals into multi-step trajectories through reasoning and error correction. The project provid

    Operates as a multimodal agent capable of interpreting screen layouts and performing multi-step tasks across desktop and mobile interfaces.

    Pythonresearch
    在 GitHub 上查看↗9,622
  • microsoft/ufomicrosoft 的头像

    microsoft/UFO

    9,017在 GitHub 上查看↗

    UFO is a multi-device task orchestrator and LLM agent orchestration framework designed to decompose natural language requests into executable task graphs. It functions as a cross-platform UI automation tool capable of performing interactions on Windows and mobile devices while routing tasks to distributed agents based on their hardware and software capabilities. The system is distinguished by its RAG-enhanced agent architecture, which integrates external documentation and previous execution traces to improve decision-making. It employs a hybrid UI detection approach that combines computer vis

    UI-focused agent for Windows operating system interaction.

    Pythonagentautomationcopilot
    在 GitHub 上查看↗9,017
  • x-plug/mobileagentX-PLUG 的头像

    X-PLUG/MobileAgent

    7,218在 GitHub 上查看↗

    MobileAgent is an LLM-powered mobile automation agent and framework designed to navigate mobile user interfaces and execute multi-step tasks. It functions as a device interface automation system that maps semantic commands to screen coordinates to perform input events across mobile operating systems. The project operates as a cross-app workflow orchestrator, switching between native on-screen interface actions and external API tools to complete sophisticated operations. It includes a visual grounding system that analyzes screenshots and interface metadata to identify elements and validate the

    Autonomous multimodal agent for mobile device interaction.

    Pythonagentandroidapp
    在 GitHub 上查看↗7,218
  • tencentqqgylab/appagentT

    TencentQQGYLab/AppAgent

    6,786在 GitHub 上查看↗

    AppAgent is an autonomous system and Android app controller that uses large language models to navigate and execute tasks within mobile applications. It functions as a mobile UI automator and element mapper, capable of performing specific application tasks by utilizing documented user interface patterns and screen navigation. The framework differentiates itself through its ability to map application navigation and generate UI documentation via autonomous exploration or human-in-the-loop demonstrations. It employs a visual-language model to process screen screenshots and UI hierarchies to dete

    Multimodal agent framework for smartphone application usage.

    Python
    在 GitHub 上查看↗6,786
  • thudm/cogvlmTHUDM 的头像

    THUDM/CogVLM

    6,742在 GitHub 上查看↗

    CogVLM is a multimodal large language model designed to integrate visual and textual data for reasoning about images and generating natural language. It functions as a visual question answering system that analyzes image content to provide detailed descriptions or answer specific questions. The project includes a visual grounding model capable of mapping text descriptions to precise bounding box coordinates within an image. It also features a vision-based automation agent that analyzes screen captures to generate execution plans and interaction coordinates for software interfaces. The system

    Provides a vision-based agent capable of analyzing screens and interacting with desktop and mobile interfaces.

    Python
    在 GitHub 上查看↗6,742
  • zai-org/cogvlmzai-org 的头像

    zai-org/CogVLM

    6,742在 GitHub 上查看↗

    CogVLM is a multimodal large language model designed for visual reasoning and multi-turn dialogue. It functions as a visual grounding model and a quantized vision model, combining text and image processing to perform complex understanding and maintain context across visual inputs. The project includes capabilities as a GUI automation agent, allowing it to analyze application screenshots, plan operational steps, and return precise screen coordinates for interface interaction. It further supports visual grounding by generating bounding box coordinates to map text descriptions to specific spatia

    Functions as an agent capable of operating and automating desktop and mobile user interfaces.

    Pythoncross-modalitylanguage-modelmulti-modal
    在 GitHub 上查看↗6,742
  1. Home
  2. Part of an Awesome List
  3. AI & Machine Learning
  4. GUI Agents