2 مستودعات
Systems that decompose graphical interfaces into structured semantic elements for machine reasoning.
Distinguishing note: Focuses on the hierarchical decomposition of visual input rather than general image processing.
Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Visual Interface Parsers. Refine with filters or upvote what's useful.
OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual screen information to execute complex, multi-step tasks across operating system environments by bridging visual interface perception with language models. Through a continuous cycle of observation and command execution, the system grounds high-level natural language instructions into precise, coordinate-based actions. The project distinguishes itself by utilizing vision-based parsing to interact with software interfaces without requiring access to underlying application progr
Decomposes complex desktop screenshots into structured semantic elements to simplify visual input for reasoning models.
Bytebot is an LLM desktop automation framework and virtual Linux desktop environment. It enables AI agents to plan and execute mouse and keyboard actions on a virtual computer using natural language, allowing for autonomous desktop automation and the integration of legacy systems that lack native APIs. The system operates as an LLM API gateway and a Model Context Protocol server, routing requests across multiple language model providers with integrated load balancing and rate limiting. It provides isolated, containerized environments where agents use visual reasoning to interpret screenshots
Interacts with user interface elements using visual intelligence to decompose graphical interfaces for machine reasoning.