awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

3 个仓库

Awesome GitHub RepositoriesGUI Task Automation

Autonomous execution of end-to-end user interface operations by identifying and interacting with screen elements.

Distinct from Automated End-to-End Testing: Unlike E2E testing, this is for general task execution and agentic operation, not software verification.

Explore 3 awesome GitHub repositories matching artificial intelligence & ml · GUI Task Automation. Refine with filters or upvote what's useful.

Awesome GUI Task Automation GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • x-plug/mobileagentX-PLUG 的头像

    X-PLUG/MobileAgent

    7,218在 GitHub 上查看↗

    MobileAgent is an LLM-powered mobile automation agent and framework designed to navigate mobile user interfaces and execute multi-step tasks. It functions as a device interface automation system that maps semantic commands to screen coordinates to perform input events across mobile operating systems. The project operates as a cross-app workflow orchestrator, switching between native on-screen interface actions and external API tools to complete sophisticated operations. It includes a visual grounding system that analyzes screenshots and interface metadata to identify elements and validate the

    Executes end-to-end operations across mobile devices by identifying interface elements and performing grounding actions.

    Pythonagentandroidapp
    在 GitHub 上查看↗7,218
  • thudm/cogvlmTHUDM 的头像

    THUDM/CogVLM

    6,742在 GitHub 上查看↗

    CogVLM is a multimodal large language model designed to integrate visual and textual data for reasoning about images and generating natural language. It functions as a visual question answering system that analyzes image content to provide detailed descriptions or answer specific questions. The project includes a visual grounding model capable of mapping text descriptions to precise bounding box coordinates within an image. It also features a vision-based automation agent that analyzes screen captures to generate execution plans and interaction coordinates for software interfaces. The system

    Generates actionable steps and coordinates to perform tasks by identifying and interacting with screen elements.

    Python
    在 GitHub 上查看↗6,742
  • zai-org/cogvlmzai-org 的头像

    zai-org/CogVLM

    6,742在 GitHub 上查看↗

    CogVLM is a multimodal large language model designed for visual reasoning and multi-turn dialogue. It functions as a visual grounding model and a quantized vision model, combining text and image processing to perform complex understanding and maintain context across visual inputs. The project includes capabilities as a GUI automation agent, allowing it to analyze application screenshots, plan operational steps, and return precise screen coordinates for interface interaction. It further supports visual grounding by generating bounding box coordinates to map text descriptions to specific spatia

    Automates end-to-end user interface operations by interpreting screenshots and interacting with screen elements.

    Pythoncross-modalitylanguage-modelmulti-modal
    在 GitHub 上查看↗6,742
  1. Home
  2. Artificial Intelligence & ML
  3. GUI Task Automation