awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

6 个仓库

Awesome GitHub RepositoriesVisual Grounding

Techniques for mapping AI-generated labels or numerical markers to specific spatial coordinates on a user interface.

Distinguishing note: None of the candidates relate to computer vision or AI-driven UI interaction; they focus on UI component libraries and design patterns.

Explore 6 awesome GitHub repositories matching artificial intelligence & ml · Visual Grounding. Refine with filters or upvote what's useful.

Awesome Visual Grounding GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • qwenlm/qwen2-vlQwenLM 的头像

    QwenLM/Qwen2-VL

    19,404在 GitHub 上查看↗

    Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw

    Maps natural language to precise spatial coordinates using 2D bounding boxes and 3D grounding.

    Jupyter Notebook
    在 GitHub 上查看↗19,404
  • othersideai/self-operating-computerOthersideAI 的头像

    OthersideAI/self-operating-computer

    10,153在 GitHub 上查看↗

    This project is a computer control framework that uses multimodal vision models to simulate mouse and keyboard inputs for automating desktop tasks. It functions as an autonomous agent and vision-based orchestrator that interprets screen visuals to interact with user interfaces. The system employs vision language models and object detection to locate and click interface elements. It utilizes visual grounding to overlay numerical markers on UI components and uses optical character recognition to map on-screen text to precise pixel coordinates. The framework supports voice-controlled computing

    The system overlays visual markers on UI components using detection models to improve AI interaction accuracy with buttons.

    Pythonautomationopenaipyautogui
    在 GitHub 上查看↗10,153
  • bytedance/ui-tarsbytedance 的头像

    bytedance/UI-TARS

    9,622在 GitHub 上查看↗

    UI-TARS is an LLM GUI automation framework and multimodal action grounding system. It functions as a GUI agent orchestrator and cross-platform device controller that uses large language models to interpret graphical interfaces and execute actions across desktop and mobile operating systems. The system translates model-generated coordinates into precise screen positions to interact with visual user interface elements. It employs a multimodal approach to interpret screen layouts and decomposes complex goals into multi-step trajectories through reasoning and error correction. The project provid

    Measures model precision by mapping coordinate outputs to specific visual elements on a screen for grounding evaluation.

    Pythonresearch
    在 GitHub 上查看↗9,622
  • apple/ml-ferretapple 的头像

    apple/ml-ferret

    8,680在 GitHub 上查看↗

    ml-ferret is a multimodal large language model framework and visual reasoning engine designed to reason about images and user interfaces. It functions as a UI grounding model and referring expression comprehension tool that maps natural language descriptions to precise pixel coordinates. The system focuses on high-resolution image analysis to identify and locate specific interface components. It employs multi-resolution image processing and region-aware visual encoding to preserve detail across different aspect ratios, enabling the model to analyze spatial relationships and functional layouts

    Maps natural language descriptions to precise pixel coordinates to identify and locate specific user interface components.

    Python
    在 GitHub 上查看↗8,680
  • x-plug/mobileagentX-PLUG 的头像

    X-PLUG/MobileAgent

    7,218在 GitHub 上查看↗

    MobileAgent is an LLM-powered mobile automation agent and framework designed to navigate mobile user interfaces and execute multi-step tasks. It functions as a device interface automation system that maps semantic commands to screen coordinates to perform input events across mobile operating systems. The project operates as a cross-app workflow orchestrator, switching between native on-screen interface actions and external API tools to complete sophisticated operations. It includes a visual grounding system that analyzes screenshots and interface metadata to identify elements and validate the

    Maps AI-generated intents to specific screen coordinates by analyzing screenshots and interface metadata.

    Pythonagentandroidapp
    在 GitHub 上查看↗7,218
  • zai-org/cogvlmzai-org 的头像

    zai-org/CogVLM

    6,742在 GitHub 上查看↗

    CogVLM is a multimodal large language model designed for visual reasoning and multi-turn dialogue. It functions as a visual grounding model and a quantized vision model, combining text and image processing to perform complex understanding and maintain context across visual inputs. The project includes capabilities as a GUI automation agent, allowing it to analyze application screenshots, plan operational steps, and return precise screen coordinates for interface interaction. It further supports visual grounding by generating bounding box coordinates to map text descriptions to specific spatia

    Generates precise bounding box coordinates to map text descriptions to specific spatial regions within an image.

    Pythoncross-modalitylanguage-modelmulti-modal
    在 GitHub 上查看↗6,742
  1. Home
  2. Artificial Intelligence & ML
  3. Visual Grounding