awesome-repositories.com
Blog
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectDespreCum realizăm clasamentulPresăServer MCP
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
microsoft avatar

microsoft/OmniParser

0
View on GitHub↗
24,377 stele·2,123 fork-uri·Jupyter Notebook·cc-by-4.0·11 vizualizări

OmniParser

OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual screen information to execute complex, multi-step tasks across operating system environments by bridging visual interface perception with language models. Through a continuous cycle of observation and command execution, the system grounds high-level natural language instructions into precise, coordinate-based actions.

The project distinguishes itself by utilizing vision-based parsing to interact with software interfaces without requiring access to underlying application programming interfaces or platform-specific accessibility frameworks. It decomposes complex screenshots into structured semantic elements and maps raw pixel data to labeled interactive components. This approach enables consistent automated workflows across varying display resolutions by normalizing coordinate spaces and relying on visual recognition rather than code-level hooks.

The software provides a comprehensive framework for autonomous agent development, allowing for the transformation of static interface captures into structured data representations. This capability facilitates accurate element identification and interaction for vision-based models during repetitive desktop tasks.

Features

  • Desktop Automation Agents - Interprets visual screen information to execute complex tasks across operating system environments through simulated user interactions.
  • Vision-Language Grounding Models - Maps natural language instructions to specific coordinate-based bounding boxes on a visual interface.
  • Agentic Orchestration Loops - Maintains a continuous cycle of screen observation and command execution to navigate through multi-step tasks.
  • Autonomous Agent Frameworks - Provides tools for building intelligent software agents capable of navigating complex graphical user interfaces.
  • Desktop Automation Frameworks - Executes complex tasks across desktop environments by combining screen parsing with vision-based language models.
  • Multimodal Interaction Engines - Bridges visual interface perception with language models to ground high-level instructions into precise coordinate-based actions.
  • Vision-Based UI Parsers - Converts visual interface screenshots into structured data representations to enable accurate element identification.
  • Visual Interface Parsers - Decomposes complex desktop screenshots into structured semantic elements to simplify visual input for reasoning models.
  • Automated Desktop Interaction Systems - Controls computer applications through visual analysis to perform repetitive tasks without direct API access.
  • Vision-Based UI Parsing Libraries - Transforms static screenshots of software interfaces into structured data formats for artificial intelligence models.
  • AI Agents - Screen parsing tool converting UI elements into structured data.
  • Computer Use - Screen parsing tool for vision-based GUI agents.
  • GUI Agents - Vision-based parsing for GUI-driven agent automation.
  • Multimodal Agents - Pure vision-based parsing for GUI agent interaction.
  • Interface Data Extraction Tools - Converts visual interface captures into structured data elements to help models ground actions accurately.
  • Cross-Platform Visual Automation Tools - Executes consistent automated workflows across different operating systems by relying on visual recognition.
  • Semantic Mapping Engines - Translates raw pixel data into labeled interactive components by matching visual features against interface primitives.

Istoric stele

Graficul istoricului de stele pentru microsoft/omniparserGraficul istoricului de stele pentru microsoft/omniparser

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Start searching with AI

Alternative open-source pentru OmniParser

Proiecte open-source similare, clasificate după numărul de funcționalități comune cu OmniParser.
  • simular-ai/agent-sAvatar simular-ai

    simular-ai/Agent-S

    11,855Vezi pe GitHub↗

    Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov

    Pythonagent-computer-interfaceai-agentscomputer-automation
    Vezi pe GitHub↗11,855
  • bytedance/ui-tars-desktopAvatar bytedance

    bytedance/UI-TARS-desktop

    36,445Vezi pe GitHub↗

    UI-TARS-desktop is a cross-platform desktop application designed to automate software interface interactions. It functions as a local agent environment that interprets graphical user interfaces through multimodal visual-language model reasoning, allowing it to navigate and manipulate software by simulating human-like mouse and keyboard inputs. The platform distinguishes itself by executing all visual recognition and decision-making logic directly on the host machine. This local inference model ensures that screen data and sensitive information remain private, as no processing is offloaded to

    TypeScriptagentagent-tarsbrowser-use
    Vezi pe GitHub↗36,445
  • bytedance/ui-tarsAvatar bytedance

    bytedance/UI-TARS

    9,622Vezi pe GitHub↗

    UI-TARS is an LLM GUI automation framework and multimodal action grounding system. It functions as a GUI agent orchestrator and cross-platform device controller that uses large language models to interpret graphical interfaces and execute actions across desktop and mobile operating systems. The system translates model-generated coordinates into precise screen positions to interact with visual user interface elements. It employs a multimodal approach to interpret screen layouts and decomposes complex goals into multi-step trajectories through reasoning and error correction. The project provid

    Pythonresearch
    Vezi pe GitHub↗9,622
  • othersideai/self-operating-computerAvatar OthersideAI

    OthersideAI/self-operating-computer

    10,153Vezi pe GitHub↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    Pythonautomationopenaipyautogui
    Vezi pe GitHub↗10,153
Vezi toate cele 30 alternative pentru OmniParser→

Întrebări frecvente

Ce face microsoft/omniparser?

OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual screen information to execute complex, multi-step tasks across operating system environments by bridging visual interface perception with language models. Through a continuous cycle of observation and command execution, the system grounds high-level natural language instructions into precise, coordinate-based actions.

Care sunt principalele funcționalități ale microsoft/omniparser?

Principalele funcționalități ale microsoft/omniparser sunt: Desktop Automation Agents, Vision-Language Grounding Models, Agentic Orchestration Loops, Autonomous Agent Frameworks, Desktop Automation Frameworks, Multimodal Interaction Engines, Vision-Based UI Parsers, Visual Interface Parsers.

Care sunt câteva alternative open-source pentru microsoft/omniparser?

Alternativele open-source pentru microsoft/omniparser includ: simular-ai/agent-s — Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through… bytedance/ui-tars-desktop — UI-TARS-desktop is a cross-platform desktop application designed to automate software interface interactions. It… bytedance/ui-tars — UI-TARS is an LLM GUI automation framework and multimodal action grounding system. It functions as a GUI agent… othersideai/self-operating-computer — This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs… microsoft/ufo — UFO is a multi-device task orchestrator and LLM agent orchestration framework designed to decompose natural language… bytebot-ai/bytebot — Bytebot is an LLM desktop automation framework and virtual Linux desktop environment. It enables AI agents to plan and…