awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
OpenGVLab avatar

OpenGVLab/ZeroGUI

0
View on GitHub↗
120 stars·9 forks·Python·Apache-2.0·13 viewsarxiv.org/abs/2505.23762↗

ZeroGUI

We propose ZeroGUI, a fully automated online reinforcement learning framework that enables GUI agents to train and adapt in interactive environments at zero human cost.

Features

  • GUI and Computer Agents - Automates online GUI learning at zero human cost.
  • Reasoning Environments - Automated online GUI learning environment.

Star history

Star history chart for opengvlab/zeroguiStar history chart for opengvlab/zerogui

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with ZeroGUI

These projects share indexed features with ZeroGUI. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • othersideai/self-operating-computerOthersideAI avatar

    OthersideAI/self-operating-computer

    10,153View on GitHub↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    View on GitHub↗10,153
  • simular-ai/agent-ssimular-ai avatar

    simular-ai/Agent-S

    11,855View on GitHub↗

    Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov

    Pythonagent-computer-interfaceai-agentscomputer-automation
    View on GitHub↗11,855
  • qwenlm/qwen2.5-vlQwenLM avatar

    QwenLM/Qwen2.5-VL

    19,480View on GitHub↗

    Qwen2.5-VL is an autoregressive multimodal transformer designed to process interleaved sequences of text and visual tokens. It integrates visual feature embeddings into a shared language model space to perform cross-modal reasoning and generate coherent responses or structured layout code. The project distinguishes itself through vision-language-action mapping, allowing it to perceive visual interfaces and translate that perception into actionable commands for operating digital screens and robotic hardware. It employs dynamic-resolution image encoding and temporal-frame video indexing to hand

    Jupyter Notebook
    View on GitHub↗19,480
  • alfworld/alfworldalfworld avatar

    alfworld/alfworld

    779View on GitHub↗

    Aligning Text and Embodied Environments for Interactive Learning Mohit Shridhar, Xingdi (Eric) Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, Matthew Hausknecht ICLR 2021

    Python
    View on GitHub↗779
Compare all 30 related projects→

Frequently asked questions

What does opengvlab/zerogui do?

We propose ZeroGUI, a fully automated online reinforcement learning framework that enables GUI agents to train and adapt in interactive environments at zero human cost.

What are the main features of opengvlab/zerogui?

The main features of opengvlab/zerogui are: GUI and Computer Agents, Reasoning Environments.

Which projects share features with opengvlab/zerogui?

Projects with overlapping indexed features include: othersideai/self-operating-computer — This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs… simular-ai/agent-s — Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through… qwenlm/qwen2.5-vl — Qwen2.5-VL is an autoregressive multimodal transformer designed to process interleaved sequences of text and visual… alfworld/alfworld — Aligning Text and Embodied Environments for Interactive Learning Mohit Shridhar, Xingdi (Eric) Yuan, Marc-Alexandre… bytedance/ui-tars-desktop — UI-TARS-desktop is a cross-platform desktop application designed to automate software interface interactions. It… bytedance/ui-tars — UI-TARS is an LLM GUI automation framework and multimodal action grounding system. It functions as a GUI agent…