4 个仓库
Architectures that process and reason across multiple data modalities to solve complex tasks.
Distinct from Multimodal Reasoning Tasks: Distinct from specific reasoning tasks or UI-based visual reasoning by being a general engine for editing tasks.
Explore 4 awesome GitHub repositories matching artificial intelligence & ml · Multimodal Reasoning Engines. Refine with filters or upvote what's useful.
ml-ferret is a multimodal large language model framework and visual reasoning engine designed to reason about images and user interfaces. It functions as a UI grounding model and referring expression comprehension tool that maps natural language descriptions to precise pixel coordinates. The system focuses on high-resolution image analysis to identify and locate specific interface components. It employs multi-resolution image processing and region-aware visual encoding to preserve detail across different aspect ratios, enabling the model to analyze spatial relationships and functional layouts
Utilizes a multimodal engine to analyze spatial relationships and functional layouts for solving complex tasks.
CogVLM is a multimodal large language model designed for visual reasoning and multi-turn dialogue. It functions as a visual grounding model and a quantized vision model, combining text and image processing to perform complex understanding and maintain context across visual inputs. The project includes capabilities as a GUI automation agent, allowing it to analyze application screenshots, plan operational steps, and return precise screen coordinates for interface interaction. It further supports visual grounding by generating bounding box coordinates to map text descriptions to specific spatia
Processes and reasons across visual and textual modalities to answer complex questions and maintain dialogues.
DeepSeek-VL 是一个多模态大型语言模型和图像到文本推理引擎。它作为一个视觉-语言模型和视觉问答系统,集成了视觉感知与语言推理,以理解和描述图像。 该项目支持多模态图像理解和文档图像分析,特别是处理网页截图和技术图表。它提供了视觉对话 AI 的功能,允许用户与视觉数据交互以提取见解,并跨不同类型的视觉信息执行复杂的推理。 该系统利用视觉-语言 Transformer 架构,将视觉 Transformer 与大型语言模型相结合。它采用多模态指令微调和投影层,将视觉特征向量与语言模型的嵌入空间对齐,以进行自回归文本生成。
Acts as a multimodal reasoning engine that extracts structured insights from diagrams and web pages.
HunyuanImage-3.0 is a diffusion-based text-to-image tool and large language model image generator designed for creating high-fidelity, photorealistic visual content. It functions as an image-to-image synthesis framework and a multimodal visual reasoning engine. The system includes a prompt refinement system that automatically rewrites sparse user inputs into detailed descriptions to improve output precision. It also employs a reasoning chain architecture to analyze image inputs and prompts, decomposing complex editing tasks into structured sub-tasks. The project covers a range of synthesis c
Implements a reasoning chain architecture that analyzes image inputs and prompts to decompose editing tasks.