awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

10 مستودعات

Awesome GitHub RepositoriesVision-Language Grounding Models

Models that map natural language instructions to specific spatial coordinates on a visual interface.

Distinguishing note: Specifically addresses the grounding of language into spatial bounding boxes.

Explore 10 awesome GitHub repositories matching artificial intelligence & ml · Vision-Language Grounding Models. Refine with filters or upvote what's useful.

Awesome Vision-Language Grounding Models GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • microsoft/omniparserالصورة الرمزية لـ microsoft

    microsoft/OmniParser

    24,377عرض على GitHub↗

    OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual screen information to execute complex, multi-step tasks across operating system environments by bridging visual interface perception with language models. Through a continuous cycle of observation and command execution, the system grounds high-level natural language instructions into precise, coordinate-based actions. The project distinguishes itself by utilizing vision-based parsing to interact with software interfaces without requiring access to underlying application progr

    Maps natural language instructions to specific coordinate-based bounding boxes on a visual interface.

    Jupyter Notebook
    عرض على GitHub↗24,377
  • zai-org/open-autoglmالصورة الرمزية لـ zai-org

    zai-org/Open-AutoGLM

    23,532عرض على GitHub↗

    Open-AutoGLM is an autonomous agent framework designed to perform complex user workflows on mobile devices. By translating natural language instructions into precise sequences of taps, scrolls, and text inputs, the system enables the automation of mobile application interactions and testing. The platform distinguishes itself through a combination of vision-language processing and reinforcement learning. It converts graphical user interfaces into structured data, allowing agents to parse screen elements and map natural language commands to coordinate-based actions. To ensure reliability, the s

    Maps natural language instructions to spatial coordinates on mobile interfaces using vision-language grounding models.

    Pythonagentphone-use-agent
    عرض على GitHub↗23,532
  • microsoft/unilmالصورة الرمزية لـ microsoft

    microsoft/unilm

    22,030عرض على GitHub↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Links text spans such as noun phrases and referring expressions to specific image regions to enable phrase grounding and comprehension.

    Pythonbeitbeit-3bitnet
    عرض على GitHub↗22,030
  • idea-research/grounded-segment-anythingالصورة الرمزية لـ IDEA-Research

    IDEA-Research/Grounded-Segment-Anything

    17,633عرض على GitHub↗

    Grounded-Segment-Anything is a suite of specialized tools for multimodal visual analysis, text-based segmentation, and generative image editing. It integrates text-to-bounding-box detection and high-precision image segmentation masks to function as a text-based image segmenter and an automated visual labeling tool. The project enables text-driven image editing by identifying objects through natural language to perform inpainting and element replacement. It further extends visual analysis into three dimensions, allowing for 3D human reconstruction and the generation of 3D bounding boxes from t

    Implements a pipeline that maps natural language prompts to spatial bounding boxes for object grounding.

    Jupyter Notebook3d-whole-body-pose-estimationautomatic-labeling-systemcaption
    عرض على GitHub↗17,633
  • simular-ai/agent-sالصورة الرمزية لـ simular-ai

    simular-ai/Agent-S

    11,855عرض على GitHub↗

    Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov

    Utilizes models that map natural language instructions to specific spatial coordinates on a visual user interface.

    Pythonagent-computer-interfaceai-agentscomputer-automation
    عرض على GitHub↗11,855
  • web-infra-dev/midsceneالصورة الرمزية لـ web-infra-dev

    web-infra-dev/midscene

    11,720عرض على GitHub↗

    Midscene is a multimodal automation framework designed to enable AI agents to perceive, navigate, and manipulate graphical user interfaces across web, mobile, and desktop environments. By leveraging vision-capable AI models, the platform interprets interface screenshots to execute tasks based on natural language instructions, removing the reliance on traditional, brittle code-based selectors. The framework distinguishes itself through its ability to decompose high-level goals into autonomous, multi-step sequences that function consistently across diverse platforms. It provides a visual ground

    Maps natural language instructions to specific screen coordinates using visual grounding.

    TypeScriptaiai-testbrowser-use
    عرض على GitHub↗11,720
  • apple/ml-ferretالصورة الرمزية لـ apple

    apple/ml-ferret

    8,680عرض على GitHub↗

    ml-ferret is a multimodal large language model framework and visual reasoning engine designed to reason about images and user interfaces. It functions as a UI grounding model and referring expression comprehension tool that maps natural language descriptions to precise pixel coordinates. The system focuses on high-resolution image analysis to identify and locate specific interface components. It employs multi-resolution image processing and region-aware visual encoding to preserve detail across different aspect ratios, enabling the model to analyze spatial relationships and functional layouts

    Maps natural language instructions to specific spatial bounding boxes on visual user interfaces.

    Python
    عرض على GitHub↗8,680
  • meta-llama/llama-modelsالصورة الرمزية لـ meta-llama

    meta-llama/llama-models

    7,643عرض على GitHub↗

    يوفر هذا المشروع إطار عمل أساسياً وتنفيذاً مرجعياً لتنفيذ نمذجة اللغة السببية والاستدلال متعدد الوسائط على الأنظمة المحلية. يتضمن مجموعة من المكونات الأساسية لإدارة أصول النماذج، وإطار عمل للضبط الدقيق، والتعريفات الهيكلية المطلوبة لإنشاء معماريات تعتمد على المحولات (Transformers). يتميز النظام بقدرته على معالجة مدخلات النصوص والصور المدمجة من خلال نماذج محولات متعددة الوسائط للاستدلال البصري وتحليل المستندات. كما يدعم نشر النماذج المكممة (Quantized)، مما يقلل من استهلاك الذاكرة من خلال تقنيات الدقة المنخفضة لتمكين الاستدلال على أجهزة الحافة (Edge devices). يغطي المشروع مجالات قدرات واسعة تشمل الضبط الدقيق الخاضع للإشراف والتكيف منخفض الرتبة (LoRA) لتخصيص النطاق، بالإضافة إلى مدير أصول شامل لتنزيل أوزان النماذج والمُرمّزات (Tokenizers) والتحقق منها وتنظيمها. تشمل الوظائف الإضافية توليد النصوص متعدد اللغات، ومعالجة السياق الطويل، والتأريض اللغوي البصري.

    Maps natural language descriptions to specific objects or spatial regions within an image.

    Python
    عرض على GitHub↗7,643
  • zai-org/cogvlmالصورة الرمزية لـ zai-org

    zai-org/CogVLM

    6,742عرض على GitHub↗

    CogVLM is a multimodal large language model designed for visual reasoning and multi-turn dialogue. It functions as a visual grounding model and a quantized vision model, combining text and image processing to perform complex understanding and maintain context across visual inputs. The project includes capabilities as a GUI automation agent, allowing it to analyze application screenshots, plan operational steps, and return precise screen coordinates for interface interaction. It further supports visual grounding by generating bounding box coordinates to map text descriptions to specific spatia

    Maps natural language instructions to specific spatial bounding boxes on a visual interface.

    Pythoncross-modalitylanguage-modelmulti-modal
    عرض على GitHub↗6,742
  • deepseek-ai/deepseek-vl2الصورة الرمزية لـ deepseek-ai

    deepseek-ai/DeepSeek-VL2

    5,302عرض على GitHub↗

    DeepSeek-VL2 هو نموذج لغوي كبير متعدد الوسائط ونظام رؤية لغوية مصمم لتحليل المشاهد المرئية وتوليد نص وصفي. يعمل كنموذج للإجابة على الأسئلة المرئية والتأريض المرئي، وقادر على استخراج المعلومات من المستندات وتحديد كائنات أو مناطق محددة داخل الصور بناءً على أوصاف نصية. يستخدم المشروع معمارية خليط من الخبراء (mixture-of-experts) لمعالجة مدخلات الصور والنصوص المدمجة. تم تحسينه للاستدلال من خلال التعبئة التزايدية (incremental prefilling)، مما يقلل من متطلبات ذاكرة GPU على الأجهزة. يغطي النموذج تحليل البيانات متعدد الوسائط وفهم المستندات المرئية، بما في ذلك تفسير المخططات والتخطيطات. يقوم بإجراء استدلال مرئي وتأريض لمطابقة الاستعلامات النصية مع المحتوى المرئي المقابل.

    Locates specific objects or regions within an image by matching them to provided textual descriptions.

    Python
    عرض على GitHub↗5,302
  1. Home
  2. Artificial Intelligence & ML
  3. Vision-Language Grounding Models