awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
bytedance avatar

bytedance/UI-TARS

0
View on GitHub↗
9,622 stars·698 forks·Python·apache-2.0·8 views

UI TARS

UI-TARS is an LLM GUI automation framework and multimodal action grounding system. It functions as a GUI agent orchestrator and cross-platform device controller that uses large language models to interpret graphical interfaces and execute actions across desktop and mobile operating systems.

The system translates model-generated coordinates into precise screen positions to interact with visual user interface elements. It employs a multimodal approach to interpret screen layouts and decomposes complex goals into multi-step trajectories through reasoning and error correction.

The project provides capabilities for cross-platform interface control, including clicking, typing, and scrolling across web, mobile, and desktop environments. It includes tools for desktop and mobile GUI interaction, automation script generation, and visual grounding evaluation to measure coordinate precision.

The framework supports hosting models on cloud platforms to provide scalable inference endpoints.

Features

  • Autonomous Agent Orchestrators - Decomposes complex goals into multi-step trajectories and manages contextual memory for autonomous interface navigation.
  • Multimodal Vision Interfaces - Integrates vision-capable models to process screenshots and text prompts for interpreting the current state of graphical user interfaces.
  • Automated Desktop Interaction Systems - Performs mouse clicks and keyboard shortcuts to navigate desktop browsers and files using visual analysis.
  • Coordinate Normalization Utilities - Translates relative model output coordinates into absolute pixel positions based on target screen resolution for visual interaction.
  • Cross-Platform GUI Controllers - Translates structured instructions into native clicks, typing, and scrolling events across web, mobile, and desktop environments.
  • Action Data Normalization - Converts generated action strings into organized data sets by scaling coordinates and normalizing formats for various model types.
  • LLM GUI Automation Frameworks - Uses large language models to interpret graphical interfaces and execute actions across desktop and mobile operating systems.
  • Multi-step Goal Execution - Decomposes complex goals into sequential trajectories using high-level reasoning and error correction.
  • Multimodal Interaction Engines - Bridges visual perception with language models to map generated coordinates to precise screen positions.
  • Text-to-Coordinate Mapping - Maps extracted screen text and visual elements to precise pixel coordinates to enable targeted automated interactions.
  • Real-Time GUI Interpretation - Processes multimodal inputs including text and images to understand and monitor dynamic user interfaces in real-time.
  • Task Planning Systems - Decomposes complex user goals into a sequence of discrete actions using a loop of reasoning and error correction.
  • GUI Agents - Operates as a multimodal agent capable of interpreting screen layouts and performing multi-step tasks across desktop and mobile interfaces.
  • GUI Action Parsing - Converts natural language model responses into machine-readable data sets for the execution of GUI interactions.
  • Cross-Platform Desktop Automation Libraries - Executes standardized interaction primitives like clicking, typing, and scrolling across desktop, mobile, and web environments.
  • Screen Space Coordinate Mappings - Converts normalized 2D model coordinates into absolute pixel positions based on current target screen dimensions.
  • Native Mobile Automation - Controls native mobile applications and emulators through automated gestures and navigation to simulate user behavior.
  • Desktop Automation - Automates workflows within a desktop environment by performing clicks, drags, and keyboard shortcuts in office software.
  • Cross-Platform Abstraction Layers - Provides a unified interface to execute standardized interaction commands across web, mobile, and desktop environments.
  • Hybrid Short-and-Long Term Memory - Maintains a hybrid system of short-term task context and long-term interaction history to improve agent decision-making.
  • Visual Grounding Execution - Maps specific coordinates to visual elements by outputting direct actions without reasoning steps to evaluate precision.
  • Visual Grounding - Measures model precision by mapping coordinate outputs to specific visual elements on a screen for grounding evaluation.
  • Model Action Runtimes - Translates language model text responses into structured data for execution within an automation runtime.
  • GUI - Converts structured action data into executable code for automating GUI tasks like clicking and dragging.
  • GUI Automation Libraries - Simulates user interactions like clicking and typing across interfaces to verify software behavior and performance.
  • GUI and Computer Agents - Automated GUI interaction using native agents.
  • Specialized Multimodal Agents - Native agent model for automated GUI interaction via screenshot perception.

Star history

Star history chart for bytedance/ui-tarsStar history chart for bytedance/ui-tars

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to UI TARS

Similar open-source projects, ranked by how many features they share with UI TARS.
  • bytedance/ui-tars-desktopbytedance avatar

    bytedance/UI-TARS-desktop

    36,445View on GitHub↗

    UI-TARS-desktop is a cross-platform desktop application designed to automate software interface interactions. It functions as a local agent environment that interprets graphical user interfaces through multimodal visual-language model reasoning, allowing it to navigate and manipulate software by simulating human-like mouse and keyboard inputs. The platform distinguishes itself by executing all visual recognition and decision-making logic directly on the host machine. This local inference model ensures that screen data and sensitive information remain private, as no processing is offloaded to

    TypeScriptagentagent-tarsbrowser-use
    View on GitHub↗36,445
  • x-plug/mobileagentX-PLUG avatar

    X-PLUG/MobileAgent

    7,218View on GitHub↗

    MobileAgent is an LLM-powered mobile automation agent and framework designed to navigate mobile user interfaces and execute multi-step tasks. It functions as a device interface automation system that maps semantic commands to screen coordinates to perform input events across mobile operating systems. The project operates as a cross-app workflow orchestrator, switching between native on-screen interface actions and external API tools to complete sophisticated operations. It includes a visual grounding system that analyzes screenshots and interface metadata to identify elements and validate the

    Pythonagentandroidapp
    View on GitHub↗7,218
  • othersideai/self-operating-computerOthersideAI avatar

    OthersideAI/self-operating-computer

    10,153View on GitHub↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    View on GitHub↗10,153
  • simular-ai/agent-ssimular-ai avatar

    simular-ai/Agent-S

    11,855View on GitHub↗

    Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov

    Pythonagent-computer-interfaceai-agentscomputer-automation
    View on GitHub↗11,855
See all 30 alternatives to UI TARS→

Frequently asked questions

What does bytedance/ui-tars do?

UI-TARS is an LLM GUI automation framework and multimodal action grounding system. It functions as a GUI agent orchestrator and cross-platform device controller that uses large language models to interpret graphical interfaces and execute actions across desktop and mobile operating systems.

What are the main features of bytedance/ui-tars?

The main features of bytedance/ui-tars are: Autonomous Agent Orchestrators, Multimodal Vision Interfaces, Automated Desktop Interaction Systems, Coordinate Normalization Utilities, Cross-Platform GUI Controllers, Action Data Normalization, LLM GUI Automation Frameworks, Multi-step Goal Execution.

What are some open-source alternatives to bytedance/ui-tars?

Open-source alternatives to bytedance/ui-tars include: bytedance/ui-tars-desktop — UI-TARS-desktop is a cross-platform desktop application designed to automate software interface interactions. It… x-plug/mobileagent — MobileAgent is an LLM-powered mobile automation agent and framework designed to navigate mobile user interfaces and… simular-ai/agent-s — Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through… othersideai/self-operating-computer — This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs… microsoft/omniparser — OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual… claude-code-best/claude-code — Claude Code is a command-line interface and multi-agent orchestration framework designed for autonomous software…