awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
zai-org avatar

zai-org/Open-AutoGLM

0
View on GitHub↗
23,532 stars·3,714 forks·Python·apache-2.0·29 viewsautoglm.z.ai/blog↗

Open AutoGLM

Open-AutoGLM is an autonomous agent framework designed to perform complex user workflows on mobile devices. By translating natural language instructions into precise sequences of taps, scrolls, and text inputs, the system enables the automation of mobile application interactions and testing.

The platform distinguishes itself through a combination of vision-language processing and reinforcement learning. It converts graphical user interfaces into structured data, allowing agents to parse screen elements and map natural language commands to coordinate-based actions. To ensure reliability, the system employs heuristic-based error recovery to navigate around interface interruptions such as pop-ups, advertisements, and network delays.

The framework provides a secure, containerized environment for executing these tasks, which isolates agent processes to protect sensitive data and maintain audit trails. Additionally, it functions as a training platform where agents refine their decision-making policies through repeated reinforcement learning cycles within virtualized mobile environments.

Features

  • UI Automation - Translates natural language commands into precise mobile UI interactions for automated testing and workflow execution.
  • Large Language Models - Leverages large language models to enable autonomous navigation and multi-step workflow execution.
  • Reinforcement Learning Environments - Provides a training platform for refining agent behaviors through reinforcement learning in virtualized mobile interfaces.
  • Mobile Testing Frameworks - Automates complex mobile user workflows by translating natural language into precise interface interactions.
  • Autonomous Agents - Develops intelligent agents capable of navigating diverse mobile interfaces through reinforcement learning.
  • Vision-Language Grounding Models - Maps natural language instructions to spatial coordinates on mobile interfaces using vision-language grounding models.
  • Cloud Sandbox Provisioning - Provides a secure, containerized cloud environment for executing and auditing automated mobile tasks.
  • Action Sequence Composers - Chains complex user input sequences into executable workflows for mobile device automation.
  • Agent Task Execution - Executes automated software tasks within secure, isolated containers with audit logging.
  • Model Feedback Loops - Implements feedback loops to iteratively refine agent decision-making policies based on task completion performance.
  • Multimodal Models - Agentic multimodal model for automated device interaction.
  • Agent Security Runtimes - Enforces security policies and isolation for agent execution environments within cloud containers.
  • Interaction Navigators - Navigates through pop-ups and network latency to maintain reliable operation chains in mobile applications.
  • Browser Automation Interfaces - Provides interfaces for programmatic interaction with mobile environments to bypass obstacles during automated testing.
  • Heuristic Selectors - Uses heuristic-based selectors to navigate dynamic mobile interfaces and recover from unexpected interruptions like pop-ups.
  • Interface Representations - Converts raw graphical user interfaces into structured data to maintain context during automated navigation.
  • Sandboxed Execution Environments - Provides isolated runtime environments for executing automated agent tasks securely within virtualized containers.

Star history

Star history chart for zai-org/open-autoglmStar history chart for zai-org/open-autoglm

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Open AutoGLM

These projects share indexed features with Open AutoGLM. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • mobile-dev-inc/maestromobile-dev-inc avatar

    mobile-dev-inc/Maestro

    10,788View on GitHub↗

    Maestro is a declarative mobile and web UI automation framework designed for end-to-end testing. It operates by querying the native accessibility tree of an application, allowing for black-box testing without requiring source code instrumentation or platform-specific dependencies. The framework distinguishes itself through a unified command syntax that abstracts interactions across Android, iOS, and web environments. It features a dynamic synchronization engine that automatically pauses test execution to account for non-deterministic animations and network-dependent content loading, ensuring

    Kotlinandroidblackbox-testingios
    View on GitHub↗10,788
  • web-infra-dev/midsceneweb-infra-dev avatar

    web-infra-dev/midscene

    11,720View on GitHub↗

    Midscene is a multimodal automation framework designed to enable AI agents to perceive, navigate, and manipulate graphical user interfaces across web, mobile, and desktop environments. By leveraging vision-capable AI models, the platform interprets interface screenshots to execute tasks based on natural language instructions, removing the reliance on traditional, brittle code-based selectors. The framework distinguishes itself through its ability to decompose high-level goals into autonomous, multi-step sequences that function consistently across diverse platforms. It provides a visual ground

    TypeScriptaiai-testbrowser-use
    View on GitHub↗11,720
  • mherrmann/heliummherrmann avatar

    mherrmann/helium

    8,306View on GitHub↗

    Helium is a Python library and high-level wrapper for Selenium designed for browser automation, functional UI testing, and web scraping. It provides a simplified interface for interacting with web applications across different browser engines. The library distinguishes itself by allowing users to identify and interact with web elements using visible text labels rather than relying exclusively on technical identifiers like XPaths or CSS selectors. This approach enables the creation of automation scripts based on human-readable labels. The toolkit covers a broad range of browser automation cap

    Python
    View on GitHub↗8,306
  • openinterpreter/openinterpreteropeninterpreter avatar

    openinterpreter/openinterpreter

    64,134View on GitHub↗

    Open Interpreter is a local language model agent framework that enables the deployment of autonomous agents capable of controlling a local operating system and its applications. It provides an execution environment where language models can run code and scripts directly on a computer to automate system tasks. The framework includes a computer control interface that allows language models to interact with web browsers and native user interfaces through programmatic commands. To ensure system stability, it utilizes a secure sandbox environment for the execution of model-generated code. The sys

    Rust
    View on GitHub↗64,134
Compare all 30 related projects→

Frequently asked questions

What does zai-org/open-autoglm do?

Open-AutoGLM is an autonomous agent framework designed to perform complex user workflows on mobile devices. By translating natural language instructions into precise sequences of taps, scrolls, and text inputs, the system enables the automation of mobile application interactions and testing.

What are the main features of zai-org/open-autoglm?

The main features of zai-org/open-autoglm are: UI Automation, Large Language Models, Reinforcement Learning Environments, Mobile Testing Frameworks, Autonomous Agents, Vision-Language Grounding Models, Cloud Sandbox Provisioning, Action Sequence Composers.

Which projects share features with zai-org/open-autoglm?

Projects with overlapping indexed features include: mobile-dev-inc/maestro — Maestro is a declarative mobile and web UI automation framework designed for end-to-end testing. It operates by… web-infra-dev/midscene — Midscene is a multimodal automation framework designed to enable AI agents to perceive, navigate, and manipulate… mherrmann/helium — Helium is a Python library and high-level wrapper for Selenium designed for browser automation, functional UI testing,… openinterpreter/openinterpreter — Open Interpreter is a local language model agent framework that enables the deployment of autonomous agents capable of… qwenlm/qwen2-vl — Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text,… livekit/livekit — LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with…