awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

22 个仓库

Awesome GitHub RepositoriesMultimodal Large Language Models

Neural architectures that process both visual and textual inputs.

Distinguishing note: Defines the core model architecture for visual-textual reasoning.

Explore 22 awesome GitHub repositories matching artificial intelligence & ml · Multimodal Large Language Models. Refine with filters or upvote what's useful.

Awesome Multimodal Large Language Models GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • anthropics/anthropic-cookbookanthropics 的头像

    anthropics/anthropic-cookbook

    45,984在 GitHub 上查看↗

    This repository is a collection of guides, notebooks, and recipes for implementing advanced prompting techniques and workflow patterns with large language models. It serves as a prompt engineering guide, an evaluation suite for scoring prompt quality, and a framework for orchestrating agents and integrating external tools. The project provides implementation patterns for building applications with Claude, specifically focusing on coordinating multiple models to split complex tasks between high-reasoning and high-efficiency agents. It includes technical demonstrations for multimodal data proce

    Ships technical demonstrations for processing visual information and parsing PDF documents using multimodal LLMs.

    Jupyter Notebook
    在 GitHub 上查看↗45,984
  • bytedance/ui-tars-desktopbytedance 的头像

    bytedance/UI-TARS-desktop

    36,445在 GitHub 上查看↗

    UI-TARS-desktop is a cross-platform desktop application designed to automate software interface interactions. It functions as a local agent environment that interprets graphical user interfaces through multimodal visual-language model reasoning, allowing it to navigate and manipulate software by simulating human-like mouse and keyboard inputs. The platform distinguishes itself by executing all visual recognition and decision-making logic directly on the host machine. This local inference model ensures that screen data and sensitive information remain private, as no processing is offloaded to

    Uses multimodal neural networks to translate visual interface elements into actionable task sequences.

    TypeScriptagentagent-tarsbrowser-use
    在 GitHub 上查看↗36,445
  • vision-cair/minigpt-4Vision-CAIR 的头像

    Vision-CAIR/MiniGPT-4

    25,679在 GitHub 上查看↗

    MiniGPT-4 is a multimodal AI framework and large language model that integrates vision encoders with language models to process and reason about combined image and text inputs. It functions as a vision-language model capable of image-based conversational AI, visual question answering, and multimodal logical reasoning. The project utilizes a pretrained vision-language integration strategy that connects a vision encoder to a language model via a linear projection layer. This approach employs frozen-backbone training to align visual representations with linguistic tokens while keeping the primar

    Implements a neural architecture that processes both visual and textual inputs for combined reasoning.

    Python
    在 GitHub 上查看↗25,679
  • openbmb/minicpm-vOpenBMB 的头像

    OpenBMB/MiniCPM-V

    25,653在 GitHub 上查看↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Functions as a large language model capable of processing text, images, and video for complex understanding.

    Python
    在 GitHub 上查看↗25,653
  • haotian-liu/llavahaotian-liu 的头像

    haotian-liu/LLaVA

    24,465在 GitHub 上查看↗

    LLaVA is a multimodal large language model architecture designed to process and interpret both image and text inputs to generate natural language responses. It functions as a research-oriented platform for visual instruction tuning, providing a framework to align language models with human intent through training on diverse datasets of paired images and text queries. The system distinguishes itself through a specialized vision-language training pipeline that connects visual data to language models using projection layers and instruction-based fine-tuning. It supports distributed inference by

    Processes both image and text inputs to generate coherent natural language responses based on visual context.

    Pythonchatbotchatgptfoundation-models
    在 GitHub 上查看↗24,465
  • openbmb/minicpm-oOpenBMB 的头像

    OpenBMB/MiniCPM-o

    23,850在 GitHub 上查看↗

    MiniCPM-o is a multimodal large language model designed to function as a real-time conversational assistant on edge devices. By mapping text, image, video, and audio inputs into a unified latent space, the system enables simultaneous cross-modal reasoning and full-duplex interaction. It is built as an edge-side inference engine, utilizing quantized model weights to maintain high-performance processing on consumer hardware. The system distinguishes itself through its integrated speech synthesis and voice cloning capabilities, which allow for the generation of expressive, personalized vocal out

    Processes real-time audio, video, and text streams using a unified vision-language model architecture.

    Pythonminicpmminicpm-vmulti-modal
    在 GitHub 上查看↗23,850
  • deepseek-ai/deepseek-ocrdeepseek-ai 的头像

    deepseek-ai/DeepSeek-OCR

    22,498在 GitHub 上查看↗

    DeepSeek-OCR is a vision processing framework designed to convert image-based text into machine-readable tokens for large language models. It functions as a document inference pipeline that encodes visual data into compact representations, enabling automated optical character recognition and document analysis workflows. The system distinguishes itself through a high-throughput architecture that utilizes hardware-accelerated batch inference to process large volumes of visual data. It incorporates dynamic resolution scaling to manage the balance between visual detail and token consumption, ensu

    Prepares visual data for ingestion into multimodal large language models.

    Python
    在 GitHub 上查看↗22,498
  • microsoft/unilmmicrosoft 的头像

    microsoft/unilm

    22,030在 GitHub 上查看↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    Integrates visual and textual data into a unified model to enable multimodal understanding and generation tasks across different input modalities.

    Pythonbeitbeit-3bitnet
    在 GitHub 上查看↗22,030
  • qwenlm/qwen2-vlQwenLM 的头像

    QwenLM/Qwen2-VL

    19,404在 GitHub 上查看↗

    Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw

    Implements a foundational neural architecture that processes both visual and textual inputs for multimodal reasoning.

    Jupyter Notebook
    在 GitHub 上查看↗19,404
  • deepseek-ai/janusdeepseek-ai 的头像

    deepseek-ai/Janus

    17,746在 GitHub 上查看↗

    Janus is a multimodal large language model and unified framework that integrates visual understanding and image generation within a single neural network. It functions as both a visual understanding model for analyzing images and a text-to-image generator. The system uses a unified transformer backbone and a multimodal latent space to bridge the gap between text and visual data. This architecture employs decoupled visual encoding and cross-modal tokenization to separate the paths for discriminative understanding and generative tasks, representing images as grids of discrete codes. The projec

    Functions as a multimodal large language model integrating visual understanding and generation.

    Pythonany-to-anyfoundation-modelsllm
    在 GitHub 上查看↗17,746
  • kornia/korniakornia 的头像

    kornia/kornia

    11,238在 GitHub 上查看↗

    Kornia is a differentiable computer vision library and cross-framework tensor vision toolset. It implements vision operations as differentiable tensors to enable integration into deep learning pipelines and supports the transpilation of operations across PyTorch, TensorFlow, JAX, and NumPy. The project provides specialized toolsets for geometric vision and stereo depth, including algorithms for 3D scene reconstruction, camera calibration, and pose estimation. It further distinguishes itself as a differentiable image augmentation framework, applying random geometric and color transformations w

    Combines computer vision operations with large language models to build applications that process both visual and textual data.

    Pythonartificial-intelligencecomputer-visiondeep-learning
    在 GitHub 上查看↗11,238
  • usagi-org/ai-goofish-monitorUsagi-org 的头像

    Usagi-org/ai-goofish-monitor

    9,002在 GitHub 上查看↗

    ai-goofish-monitor is an AI-driven marketplace monitor and containerized web scraper designed to track online listings. It uses multimodal large language models and natural language prompts to analyze product text and images, determining if items meet specific requirements. The system employs an anti-detection workflow that rotates network proxies and authenticated accounts to bypass rate limits. It captures browser cookies and session states to mimic real user behavior during automated requests. The project includes a task scheduler using cron expressions and an embedded SQLite database for

    Employs multimodal large language models to process both visual and textual product data.

    Pythonaiplaywright
    在 GitHub 上查看↗9,002
  • apple/ml-ferretapple 的头像

    apple/ml-ferret

    8,680在 GitHub 上查看↗

    ml-ferret is a multimodal large language model framework and visual reasoning engine designed to reason about images and user interfaces. It functions as a UI grounding model and referring expression comprehension tool that maps natural language descriptions to precise pixel coordinates. The system focuses on high-resolution image analysis to identify and locate specific interface components. It employs multi-resolution image processing and region-aware visual encoding to preserve detail across different aspect ratios, enabling the model to analyze spatial relationships and functional layouts

    Provides a neural architecture that integrates visual encoders with language models for multimodal reasoning over images and text.

    Python
    在 GitHub 上查看↗8,680
  • zai-org/glm-4zai-org 的头像

    zai-org/GLM-4

    7,058在 GitHub 上查看↗

    GLM-4 is a large language model and fine-tuning framework designed for human-like text production, complex reasoning, and multilingual conversation. It functions as a multimodal system capable of processing high-resolution visual content and as a long-context model designed to analyze documents with a context window of up to one million tokens. The project differentiates itself through a function calling interface that enables AI agent development by connecting the model to external APIs and real-time web browsing. It includes specialized capabilities for generating functional programming cod

    Implements a neural architecture capable of processing and reasoning over both high-resolution visual content and text.

    Pythonchatglmchatglm-6bglm
    在 GitHub 上查看↗7,058
  • tencentqqgylab/appagentT

    TencentQQGYLab/AppAgent

    6,786在 GitHub 上查看↗

    AppAgent is an autonomous system and Android app controller that uses large language models to navigate and execute tasks within mobile applications. It functions as a mobile UI automator and element mapper, capable of performing specific application tasks by utilizing documented user interface patterns and screen navigation. The framework differentiates itself through its ability to map application navigation and generate UI documentation via autonomous exploration or human-in-the-loop demonstrations. It employs a visual-language model to process screen screenshots and UI hierarchies to dete

    Integrates multimodal models to process screen screenshots and UI hierarchies for determining interaction steps.

    Python
    在 GitHub 上查看↗6,786
  • thudm/cogvlmTHUDM 的头像

    THUDM/CogVLM

    6,742在 GitHub 上查看↗

    CogVLM is a multimodal large language model designed to integrate visual and textual data for reasoning about images and generating natural language. It functions as a visual question answering system that analyzes image content to provide detailed descriptions or answer specific questions. The project includes a visual grounding model capable of mapping text descriptions to precise bounding box coordinates within an image. It also features a vision-based automation agent that analyzes screen captures to generate execution plans and interaction coordinates for software interfaces. The system

    Implements a neural architecture that processes both visual and textual inputs for complex reasoning.

    Python
    在 GitHub 上查看↗6,742
  • zai-org/cogvlmzai-org 的头像

    zai-org/CogVLM

    6,742在 GitHub 上查看↗

    CogVLM is a multimodal large language model designed for visual reasoning and multi-turn dialogue. It functions as a visual grounding model and a quantized vision model, combining text and image processing to perform complex understanding and maintain context across visual inputs. The project includes capabilities as a GUI automation agent, allowing it to analyze application screenshots, plan operational steps, and return precise screen coordinates for interface interaction. It further supports visual grounding by generating bounding box coordinates to map text descriptions to specific spatia

    Implements a large-scale neural architecture that processes both visual and textual inputs for reasoning.

    Pythoncross-modalitylanguage-modelmulti-modal
    在 GitHub 上查看↗6,742
  • qwenlm/qwen-vlQwenLM 的头像

    QwenLM/Qwen-VL

    6,535在 GitHub 上查看↗

    Processes images and text together to generate responses for visual question answering and captioning.

    Pythonlarge-language-modelsvision-language-model
    在 GitHub 上查看↗6,535
  • deepseek-ai/deepseek-vl2deepseek-ai 的头像

    deepseek-ai/DeepSeek-VL2

    5,302在 GitHub 上查看↗

    DeepSeek-VL2 是一个多模态大语言模型和视觉语言系统,旨在分析视觉场景并生成描述性文本。它作为一个视觉问答和视觉定位模型,能够从文档中提取信息,并根据文本描述定位图像中的特定对象或区域。 该项目利用专家混合(mixture-of-experts)架构来处理组合的图像和文本输入。它通过增量预填充(incremental prefilling)针对推理进行了优化,从而降低了硬件上的 GPU 内存需求。 该模型涵盖多模态数据分析和视觉文档理解,包括对图表和布局的解释。它执行视觉推理和定位,以将文本查询与相应的视觉内容进行匹配。

    Implements a neural architecture capable of processing both visual and textual inputs for reasoning.

    Python
    在 GitHub 上查看↗5,302
  • zai-org/glm-4.5zai-org 的头像

    zai-org/GLM-4.5

    4,210在 GitHub 上查看↗

    GLM-4.5 is a multimodal large language model and advanced reasoning system. It functions as an AI coding assistant, an autonomous AI agent, and a multimodal content generator capable of processing and generating text, images, audio, and video within a single unified system. The project is distinguished by its deep reasoning capabilities, utilizing chain-of-thought processing to solve complex mathematical, logical, and technical problems. It features an agentic architecture that allows for autonomous task execution, long-horizon goal planning, and the ability to interact with external tools an

    Processes and generates text, images, audio, and video within a single unified system.

    Pythonagentglmllm
    在 GitHub 上查看↗4,210
上一个12下一个
  1. Home
  2. Artificial Intelligence & ML
  3. Multimodal Large Language Models

探索子标签

  • FrameworksComprehensive toolkits for instantiating and managing transformer-based model architectures. **Distinct from Multimodal Large Language Models:** Distinct from Multimodal Large Language Models: focuses on the framework infrastructure for model instantiation rather than the model architecture itself.
  • Implementation PatternsTechnical demonstrations and blueprints for applying multimodal models to real-world data tasks. **Distinct from Multimodal Large Language Models:** Focuses on the application patterns and demonstrations rather than the underlying model architecture