awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 مستودعات

Awesome GitHub RepositoriesVisual Interface Parsers

Systems that decompose graphical interfaces into structured semantic elements for machine reasoning.

Distinguishing note: Focuses on the hierarchical decomposition of visual input rather than general image processing.

Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Visual Interface Parsers. Refine with filters or upvote what's useful.

Awesome Visual Interface Parsers GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • microsoft/omniparserالصورة الرمزية لـ microsoft

    microsoft/OmniParser

    24,377عرض على GitHub↗

    OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual screen information to execute complex, multi-step tasks across operating system environments by bridging visual interface perception with language models. Through a continuous cycle of observation and command execution, the system grounds high-level natural language instructions into precise, coordinate-based actions. The project distinguishes itself by utilizing vision-based parsing to interact with software interfaces without requiring access to underlying application progr

    Decomposes complex desktop screenshots into structured semantic elements to simplify visual input for reasoning models.

    Jupyter Notebook
    عرض على GitHub↗24,377
  • bytebot-ai/bytebotالصورة الرمزية لـ bytebot-ai

    bytebot-ai/bytebot

    10,413عرض على GitHub↗

    Bytebot is an LLM desktop automation framework and virtual Linux desktop environment. It enables AI agents to plan and execute mouse and keyboard actions on a virtual computer using natural language, allowing for autonomous desktop automation and the integration of legacy systems that lack native APIs. The system operates as an LLM API gateway and a Model Context Protocol server, routing requests across multiple language model providers with integrated load balancing and rate limiting. It provides isolated, containerized environments where agents use visual reasoning to interpret screenshots

    Interacts with user interface elements using visual intelligence to decompose graphical interfaces for machine reasoning.

    TypeScriptagentagentic-aiagents
    عرض على GitHub↗10,413
  1. Home
  2. Artificial Intelligence & ML
  3. Visual Interface Parsers