awesome-repositories.com
Blog
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetÀ proposNotre méthodologiePresseServeur MCP
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
OthersideAI avatar

OthersideAI/self-operating-computer

0
View on GitHub↗
10,153 stars·1,408 forks·Python·mit·5 vueswww.hyperwriteai.com/self-operating-computer↗

Self Operating Computer

This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces.

The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates.

The framework supports voice-controlled computing by translating spoken commands into text-based objectives. It manages a full automation loop encompassing state observation through screenshots, action planning via cloud or local APIs, and the execution of synthetic inputs.

Features

  • GUI and Computer Agents - Implements an autonomous agent that uses vision models to interact with operating system GUIs and automate desktop tasks.
  • Computer Automation Interfaces - Provides a control layer that simulates human input and uses visual analysis to automate desktop tasks.
  • Autonomous UI Interaction - Interacts with software applications by mapping on-screen text and visual elements to precise clickable coordinates.
  • Multimodal AI Orchestrators - Coordinates vision and language models to simulate mouse and keyboard actions for unified agentic workflows.
  • Multimodal Vision Interfaces - Provides the integration layer for processing screen images through multimodal AI models for high-level action planning.
  • Visual UI Labeling - Overlays numerical markers on detected UI components to help the AI reference specific elements by ID.
  • Screen Text Extractors - Uses OCR on arbitrary screen regions to map clickable text and buttons to screen coordinates.
  • Text-to-Coordinate Mapping - Generates a coordinate map of on-screen text using OCR to allow precise clicking of specific elements.
  • Visual Grounding - The system overlays visual markers on UI components using detection models to improve AI interaction accuracy with buttons.
  • Multimodal Agents - Implements an autonomous agent capable of reasoning across vision and language to perform actions.
  • Virtual Input Simulation - Simulates programmatic mouse movements and keyboard strokes to automate operating system interactions.
  • Multimodal Desktop Observers - Combines visual screenshots and multimodal models to monitor and determine the current state of the desktop environment.
  • Computer Vision Screen Interaction Tools - Locates visual elements on a display through multimodal vision models to execute automated interactions.
  • OCR Coordinate Mapping - Uses optical character recognition to translate on-screen text labels into precise X and Y pixel coordinates.
  • Visual Localization Tools - Identifies and localizes UI components through visual analysis and coordinate mapping.
  • AI Model Integrations - Provides adapters and interfaces for connecting to both cloud-based and local vision models to drive interactions.
  • Voice Controlled Computing - Executes complex computer tasks and system objectives through spoken commands captured by audio hardware.
  • Voice-Controlled Goal Definition - Translates spoken user commands into text-based objectives to seed the autonomous agent loop.
  • Speech-to-Text Pipelines - Implements automated workflows that convert spoken audio input into text-based goals for the agent loop.
  • Visual Element Identification - Identifies interface components by searching for specific images or patterns within screen captures.
  • AI Agents - Experimental framework for AI-driven computer operation.
  • Autonomous Agent Frameworks - Enables multimodal models to control computer interfaces autonomously.
  • Computer Use - Multimodal framework for operating desktop applications.
  • Personal AI Assistants - Automates repetitive desktop and browser tasks via human-like interaction.

Historique des stars

Graphique de l'historique des stars pour othersideai/self-operating-computerGraphique de l'historique des stars pour othersideai/self-operating-computer

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Alternatives open source à Self Operating Computer

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec Self Operating Computer.
  • simular-ai/agent-sAvatar de simular-ai

    simular-ai/Agent-S

    11,855Voir sur GitHub↗

    Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov

    Pythonagent-computer-interfaceai-agentscomputer-automation
    Voir sur GitHub↗11,855
  • bytedance/ui-tarsAvatar de bytedance

    bytedance/UI-TARS

    9,622Voir sur GitHub↗

    UI-TARS is an LLM GUI automation framework and multimodal action grounding system. It functions as a GUI agent orchestrator and cross-platform device controller that uses large language models to interpret graphical interfaces and execute actions across desktop and mobile operating systems. The system translates model-generated coordinates into precise screen positions to interact with visual user interface elements. It employs a multimodal approach to interpret screen layouts and decomposes complex goals into multi-step trajectories through reasoning and error correction. The project provid

    Pythonresearch
    Voir sur GitHub↗9,622
  • bytedance/ui-tars-desktopAvatar de bytedance

    bytedance/UI-TARS-desktop

    36,445Voir sur GitHub↗

    UI-TARS-desktop is a cross-platform desktop application designed to automate software interface interactions. It functions as a local agent environment that interprets graphical user interfaces through multimodal visual-language model reasoning, allowing it to navigate and manipulate software by simulating human-like mouse and keyboard inputs. The platform distinguishes itself by executing all visual recognition and decision-making logic directly on the host machine. This local inference model ensures that screen data and sensitive information remain private, as no processing is offloaded to

    TypeScriptagentagent-tarsbrowser-use
    Voir sur GitHub↗36,445
  • openinterpreter/open-interpreterAvatar de openinterpreter

    openinterpreter/open-interpreter

    63,998Voir sur GitHub↗

    Open Interpreter is an autonomous agent runtime that translates natural language instructions into executable code to interact with local software and operating systems. It functions as an orchestration framework that connects language models to a secure execution environment, enabling the development of agents capable of managing system resources and performing complex tasks. To ensure safety, the system mandates explicit user verification before executing any generated code and provides robust isolation through containerized sandboxing. The project distinguishes itself through its deep inte

    Rustchatgptgpt-4interpreter
    Voir sur GitHub↗63,998
Voir les 30 alternatives à Self Operating Computer→

Questions fréquentes

Que fait othersideai/self-operating-computer ?

This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces.

Quelles sont les fonctionnalités principales de othersideai/self-operating-computer ?

Les fonctionnalités principales de othersideai/self-operating-computer sont : GUI and Computer Agents, Computer Automation Interfaces, Autonomous UI Interaction, Multimodal AI Orchestrators, Multimodal Vision Interfaces, Visual UI Labeling, Screen Text Extractors, Text-to-Coordinate Mapping.

Quelles sont les alternatives open-source à othersideai/self-operating-computer ?

Les alternatives open-source à othersideai/self-operating-computer incluent : simular-ai/agent-s — Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through… bytedance/ui-tars — UI-TARS is an LLM GUI automation framework and multimodal action grounding system. It functions as a GUI agent… bytedance/ui-tars-desktop — UI-TARS-desktop is a cross-platform desktop application designed to automate software interface interactions. It… openinterpreter/open-interpreter — Open Interpreter is an autonomous agent runtime that translates natural language instructions into executable code to… microsoft/omniparser — OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual… lmeszinc/azurlaneautoscript — AzurLaneAutoScript is a mobile game automation system designed to perform repetitive gameplay tasks unattended. It…