awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

5 个仓库

Awesome GitHub RepositoriesMultimodal Visual Understanding

Systems that integrate visual and language data to perform complex reasoning and maintain context across visual inputs.

Distinct from Multimodal Understanding: The existing candidates focus on document-specific or audio-specific understanding rather than general multimodal visual reasoning.

Explore 5 awesome GitHub repositories matching artificial intelligence & ml · Multimodal Visual Understanding. Refine with filters or upvote what's useful.

Awesome Multimodal Visual Understanding GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • zai-org/cogvlmzai-org 的头像

    zai-org/CogVLM

    6,742在 GitHub 上查看↗

    CogVLM is a multimodal large language model designed for visual reasoning and multi-turn dialogue. It functions as a visual grounding model and a quantized vision model, combining text and image processing to perform complex understanding and maintain context across visual inputs. The project includes capabilities as a GUI automation agent, allowing it to analyze application screenshots, plan operational steps, and return precise screen coordinates for interface interaction. It further supports visual grounding by generating bounding box coordinates to map text descriptions to specific spatia

    Integrates visual and language data to perform complex understanding and maintain context across visual inputs.

    Pythoncross-modalitylanguage-modelmulti-modal
    在 GitHub 上查看↗6,742
  • deepseek-ai/deepseek-vl2deepseek-ai 的头像

    deepseek-ai/DeepSeek-VL2

    5,302在 GitHub 上查看↗

    DeepSeek-VL2 是一个多模态大语言模型和视觉语言系统,旨在分析视觉场景并生成描述性文本。它作为一个视觉问答和视觉定位模型,能够从文档中提取信息,并根据文本描述定位图像中的特定对象或区域。 该项目利用专家混合(mixture-of-experts)架构来处理组合的图像和文本输入。它通过增量预填充(incremental prefilling)针对推理进行了优化,从而降低了硬件上的 GPU 内存需求。 该模型涵盖多模态数据分析和视觉文档理解,包括对图表和布局的解释。它执行视觉推理和定位,以将文本查询与相应的视觉内容进行匹配。

    Processes combined image and text inputs to perform complex multimodal visual reasoning.

    Python
    在 GitHub 上查看↗5,302
  • openimages/datasetopenimages 的头像

    openimages/dataset

    4,366在 GitHub 上查看↗

    该项目是一个计算机视觉数据集和图像标注仓库,专为训练和评估机器学习模型而设计。它提供了一个大型标注图像集合,作为目标检测基准和像素级分割数据源。 该仓库作为多模态视觉数据集脱颖而出,通过将图像与同步的语音、文本和鼠标轨迹配对,支持叙事理解。它还通过包含人口统计属性和详尽的标注,支持模型公平性分析。 该数据集涵盖了广泛的计算机视觉能力,包括通过边界框进行的目标检测、使用像素掩码的图像实例分割,以及通过对象-属性三元组进行的视觉关系映射。它还支持点级分类、分层文本识别,以及基于类或属性过滤检索精选数据集子集。

    Integrates visual and language data, linking voice traces and narratives to image regions for complex reasoning.

    Python
    在 GitHub 上查看↗4,366
  • deepseek-ai/deepseek-vldeepseek-ai 的头像

    deepseek-ai/DeepSeek-VL

    4,134在 GitHub 上查看↗

    DeepSeek-VL 是一个多模态大型语言模型和图像到文本推理引擎。它作为一个视觉-语言模型和视觉问答系统,集成了视觉感知与语言推理,以理解和描述图像。 该项目支持多模态图像理解和文档图像分析,特别是处理网页截图和技术图表。它提供了视觉对话 AI 的功能,允许用户与视觉数据交互以提取见解,并跨不同类型的视觉信息执行复杂的推理。 该系统利用视觉-语言 Transformer 架构,将视觉 Transformer 与大型语言模型相结合。它采用多模态指令微调和投影层,将视觉特征向量与语言模型的嵌入空间对齐,以进行自回归文本生成。

    Integrates visual and language data to analyze images and diagrams through complex multimodal reasoning.

    Python
    在 GitHub 上查看↗4,134
  • moonshotai/kimi-codeMoonshotAI 的头像

    MoonshotAI/kimi-code

    2,473在 GitHub 上查看↗

    Kimi-code is a command-line interface and orchestration framework designed to integrate autonomous AI agents into software development workflows. It functions as a terminal-based assistant that manages multi-step coding tasks, including planning, file system modifications, shell command execution, and test running, all while maintaining conversational context within a local development environment. The project distinguishes itself through a focus on secure, autonomous agent orchestration and granular control over AI interactions. It enforces strict security by requiring explicit user approval

    Analyzes screen recordings and video clips alongside text prompts to understand visual context.

    TypeScript
    在 GitHub 上查看↗2,473
  1. Home
  2. Artificial Intelligence & ML
  3. Multimodal Visual Understanding