10 个仓库
Models that map natural language instructions to specific spatial coordinates on a visual interface.
Distinguishing note: Specifically addresses the grounding of language into spatial bounding boxes.
Explore 10 awesome GitHub repositories matching artificial intelligence & ml · Vision-Language Grounding Models. Refine with filters or upvote what's useful.
OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual screen information to execute complex, multi-step tasks across operating system environments by bridging visual interface perception with language models. Through a continuous cycle of observation and command execution, the system grounds high-level natural language instructions into precise, coordinate-based actions. The project distinguishes itself by utilizing vision-based parsing to interact with software interfaces without requiring access to underlying application progr
Maps natural language instructions to specific coordinate-based bounding boxes on a visual interface.
Open-AutoGLM is an autonomous agent framework designed to perform complex user workflows on mobile devices. By translating natural language instructions into precise sequences of taps, scrolls, and text inputs, the system enables the automation of mobile application interactions and testing. The platform distinguishes itself through a combination of vision-language processing and reinforcement learning. It converts graphical user interfaces into structured data, allowing agents to parse screen elements and map natural language commands to coordinate-based actions. To ensure reliability, the s
Maps natural language instructions to spatial coordinates on mobile interfaces using vision-language grounding models.
This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec
Links text spans such as noun phrases and referring expressions to specific image regions to enable phrase grounding and comprehension.
Grounded-Segment-Anything is a suite of specialized tools for multimodal visual analysis, text-based segmentation, and generative image editing. It integrates text-to-bounding-box detection and high-precision image segmentation masks to function as a text-based image segmenter and an automated visual labeling tool. The project enables text-driven image editing by identifying objects through natural language to perform inpainting and element replacement. It further extends visual analysis into three dimensions, allowing for 3D human reconstruction and the generation of 3D bounding boxes from t
Implements a pipeline that maps natural language prompts to spatial bounding boxes for object grounding.
Agent-S is a multimodal AI agent and LLM desktop automation framework designed to control operating systems through graphical user interface interactions. It functions as a computer use interface, utilizing vision-language grounding to translate natural language goals into precise screen coordinates and system actions. The project differentiates itself by combining structured accessibility tree inspection with vision-based element localization. It manages cross-application workflows by mapping conceptual descriptions to physical pixels and simulating low-level keyboard and mouse events to mov
Utilizes models that map natural language instructions to specific spatial coordinates on a visual user interface.
Midscene is a multimodal automation framework designed to enable AI agents to perceive, navigate, and manipulate graphical user interfaces across web, mobile, and desktop environments. By leveraging vision-capable AI models, the platform interprets interface screenshots to execute tasks based on natural language instructions, removing the reliance on traditional, brittle code-based selectors. The framework distinguishes itself through its ability to decompose high-level goals into autonomous, multi-step sequences that function consistently across diverse platforms. It provides a visual ground
Maps natural language instructions to specific screen coordinates using visual grounding.
ml-ferret is a multimodal large language model framework and visual reasoning engine designed to reason about images and user interfaces. It functions as a UI grounding model and referring expression comprehension tool that maps natural language descriptions to precise pixel coordinates. The system focuses on high-resolution image analysis to identify and locate specific interface components. It employs multi-resolution image processing and region-aware visual encoding to preserve detail across different aspect ratios, enabling the model to analyze spatial relationships and functional layouts
Maps natural language instructions to specific spatial bounding boxes on visual user interfaces.
本项目提供了一个基础框架和参考实现,用于在本地系统上执行因果语言建模和多模态推理。它包含用于管理模型资产的核心组件、微调框架以及实例化 Transformer 架构所需的结构定义。 该系统的特点是能够通过多模态 Transformer 模型处理文本和图像组合输入,以进行视觉推理和文档分析。它还支持部署量化模型,通过低精度技术减少内存占用,从而在边缘设备上实现推理。 该项目涵盖了广泛的能力领域,包括用于领域定制的监督微调和低秩自适应 (LoRA),以及用于下载、验证和组织模型权重与分词器的综合资产管理器。其他功能还包括多语言文本生成、长上下文处理和视觉语言接地。
Maps natural language descriptions to specific objects or spatial regions within an image.
CogVLM is a multimodal large language model designed for visual reasoning and multi-turn dialogue. It functions as a visual grounding model and a quantized vision model, combining text and image processing to perform complex understanding and maintain context across visual inputs. The project includes capabilities as a GUI automation agent, allowing it to analyze application screenshots, plan operational steps, and return precise screen coordinates for interface interaction. It further supports visual grounding by generating bounding box coordinates to map text descriptions to specific spatia
Maps natural language instructions to specific spatial bounding boxes on a visual interface.
DeepSeek-VL2 是一个多模态大语言模型和视觉语言系统,旨在分析视觉场景并生成描述性文本。它作为一个视觉问答和视觉定位模型,能够从文档中提取信息,并根据文本描述定位图像中的特定对象或区域。 该项目利用专家混合(mixture-of-experts)架构来处理组合的图像和文本输入。它通过增量预填充(incremental prefilling)针对推理进行了优化,从而降低了硬件上的 GPU 内存需求。 该模型涵盖多模态数据分析和视觉文档理解,包括对图表和布局的解释。它执行视觉推理和定位,以将文本查询与相应的视觉内容进行匹配。
Locates specific objects or regions within an image by matching them to provided textual descriptions.