awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

开源视觉语言模型

排名更新于 2026年6月30日

For 开源视觉语言模型, the strongest matches are salesforce/lavis (LAVIS is a comprehensive framework and library of pre-trained), thudm/glm-4 (GLM-4 is an open-weights multimodal language model that processes) and vision-cair/minigpt-4 (MiniGPT-4 is an open-source multimodal vision-language model that processes). salesforce/blip and vikhyat/moondream round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

探索专为高级图像理解和多模态视觉推理任务而设计的开源模型与框架。

开源视觉语言模型

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • salesforce/lavissalesforce 的头像

    salesforce/LAVIS

    11,236在 GitHub 上查看↗

    LAVIS is a multimodal large language model framework and vision-language model library. It provides tools for training and evaluating models that integrate visual, textual, and audio data, serving as a cross-modal feature extractor and a zero-shot visual reasoning engine. The framework distinguishes itself by using frozen-backbone integration, where pretrained encoders remain non-trainable while lightweight adapter layers are updated. It employs cross-modal feature alignment to map different representations into a shared embedding space and utilizes a modular model wrapper to swap vision and

    LAVIS is a comprehensive framework and library of pre-trained vision-language models (like BLIP and ALBEF) that directly supports image captioning, visual question answering, and zero-shot multimodal reasoning, making it a flagship open-source toolbox for your needs.

    Jupyter NotebookVision-Language ModelsVisual Question AnsweringZero-Shot Inference Engines
    在 GitHub 上查看↗11,236
  • thudm/glm-4THUDM 的头像

    THUDM/GLM-4

    7,059在 GitHub 上查看↗

    GLM-4 is an open weights large language model designed as a multimodal chat system. It functions as a reasoning-focused and multilingual model capable of processing and generating responses across text and visual data types. The model is distinguished by its function-calling capabilities, allowing it to interface with external tools and APIs to execute tasks and retrieve real-time information. It is optimized for complex logical reasoning, mathematical problem solving, and deep research involving long-form content generation. Broad capabilities include multilingual text generation, the creat

    GLM-4 is an open-weights multimodal language model that processes both text and images, supports fine-tuning and function calling, and is built for reasoning and visual understanding, directly matching the need for an open-source GPT-4 Vision alternative.

    PythonModel Fine-TuningOpen-Weights ModelsMultimodal Models
    在 GitHub 上查看↗7,059
  • vision-cair/minigpt-4Vision-CAIR 的头像

    Vision-CAIR/MiniGPT-4

    25,679在 GitHub 上查看↗

    MiniGPT-4 is a multimodal AI framework and large language model that integrates vision encoders with language models to process and reason about combined image and text inputs. It functions as a vision-language model capable of image-based conversational AI, visual question answering, and multimodal logical reasoning. The project utilizes a pretrained vision-language integration strategy that connects a vision encoder to a language model via a linear projection layer. This approach employs frozen-backbone training to align visual representations with linguistic tokens while keeping the primar

    MiniGPT-4 is an open-source multimodal vision-language model that processes text and images together, supporting visual question answering, image captioning, and conversational reasoning, with pre-trained weights and fine-tuning capability — exactly the kind of model this search is after.

    PythonMultimodal Fine-TuningVision-Language Models
    在 GitHub 上查看↗25,679
  • salesforce/blipsalesforce 的头像

    salesforce/BLIP

    5,676在 GitHub 上查看↗

    BLIP is a vision-language model framework that combines contrastive, matching, and language modeling objectives to align images with text. Built on a multimodal encoder-decoder architecture, it supports distributed data-parallel training with cosine learning rate scheduling and sliding-window metric tracking for training stability. The framework provides capabilities for image captioning, visual question answering, and cross-modal retrieval, scoring semantic alignment between images and text through learned embeddings. It includes toolkits for fine-tuning pre-trained models on custom datasets

    BLIP is a multimodal vision-language model framework that handles image captioning, visual question answering, and cross-modal retrieval, with open-source weights, fine-tuning support, and pre-training on large datasets — exactly the kind of image-understanding AI you’re looking for.

    Jupyter NotebookImage CaptioningVisual Question Answering
    在 GitHub 上查看↗5,676
  • vikhyat/moondreamvikhyat 的头像

    vikhyat/moondream

    9,769在 GitHub 上查看↗

    Moondream is a small-scale vision language model designed to reason across images to generate captions and answer natural language questions. It functions as an edge-optimized system capable of performing visual question answering, image captioning, and object detection. The project distinguishes itself through a lightweight architecture designed for local inference on embedded devices, workstations, and air-gapped hardware. It supports the execution of models on local GPUs and Apple Silicon to ensure data privacy and low latency. The system's capabilities include identifying precise object

    Moondream is an open-source vision-language model that generates captions and answers questions about images, supporting local inference and fine-tuning — exactly the multimodal image understanding tool you are looking for.

    PythonImage CaptioningMultimodal Fine-TuningVision-Language Models
    在 GitHub 上查看↗9,769
  • clovaai/donutclovaai 的头像

    clovaai/donut

    6,789在 GitHub 上查看↗

    Donut is an OCR-free document transformer and end-to-end document parser. It functions as a neural network that converts unstructured document images directly into structured data or text without the use of an external optical character recognition engine. The project includes a synthetic document generator to create artificial images and ground-truth labels for training. It employs a transformer model to perform visual question answering and document image classification based on visual layout and text. The system covers several document understanding capabilities, including structured info

    Donut is a multimodal vision-language model specialized in document understanding — it supports visual question answering and captioning on document images, but its domain focus means it may not handle general images like GPT-4 Vision.

    PythonImage-to-Text TransformersModel Fine-TuningMultimodal Fine-Tuning
    在 GitHub 上查看↗6,789
  • opengvlab/internvlOpenGVLab 的头像

    OpenGVLab/InternVL

    10,061在 GitHub 上查看↗

    InternVL is a vision-language model framework that fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning. It provides a system for multimodal inference and dialogue, enabling the processing of images and text to answer questions or generate descriptions. The project is distinguished by its high-resolution image processing, which uses dynamic tiling to maintain detail for images up to 4K resolution, and its chain-of-thought visual reasoning for solving complex mathematical and spatial problems. It also supports temporal frame sampling

    InternVL is an open-source vision-language model that fuses a visual encoder with a large language model for multimodal reasoning, covering image captioning, visual question answering, and high-resolution image processing, directly matching the search for an open-source alternative to GPT-4 Vision.

    PythonImage CaptioningModel Fine-TuningVision-Language Models
    在 GitHub 上查看↗10,061
  • haotian-liu/llavahaotian-liu 的头像

    haotian-liu/LLaVA

    24,465在 GitHub 上查看↗

    LLaVA is a multimodal large language model architecture designed to process and interpret both image and text inputs to generate natural language responses. It functions as a research-oriented platform for visual instruction tuning, providing a framework to align language models with human intent through training on diverse datasets of paired images and text queries. The system distinguishes itself through a specialized vision-language training pipeline that connects visual data to language models using projection layers and instruction-based fine-tuning. It supports distributed inference by

    LLaVA is an open-source multimodal vision-language model that processes both image and text inputs to generate natural language responses, supporting image captioning, visual question answering, fine-tuning, and zero-shot inference — squarely matching the need for a GPT-4 Vision alternative.

    PythonModel Fine-Tuning
    在 GitHub 上查看↗24,465
  • mlfoundations/open_clipmlfoundations 的头像

    mlfoundations/open_clip

    13,935在 GitHub 上查看↗

    Open CLIP is an open source framework for training and deploying Contrastive Language-Image Pre-training models. It serves as a vision-language training framework and multimodal embedding engine that maps images and text into a shared vector space for similarity searches and zero-shot classification. The project provides a toolkit for distributed training of contrastive models and includes an image-to-text generative model for producing natural language descriptions. It supports custom text encoder integration and utilizes teacher-student model distillation to transfer knowledge from large pr

    Open CLIP is a framework for training and deploying contrastive vision-language models like CLIP, a foundational open-source multimodal model that maps images and text into a shared space for zero-shot classification and captioning — it covers image understanding, fine-tuning, and pre-trained weights, though it is more of an embedding engine than a generative VQA system like GPT-4V.

    PythonContrastive Pre-trainingImage CaptioningImage-to-Text Transformers
    在 GitHub 上查看↗13,935
  • karpathy/neuraltalk2karpathy 的头像

    karpathy/neuraltalk2

    5,588在 GitHub 上查看↗

    Neuraltalk2 is a deep learning vision system designed for automatic image captioning. Built with PyTorch, it utilizes a hybrid architecture that combines a convolutional neural network encoder with a recurrent neural network decoder to generate textual descriptions from visual input. The project features a GPU-accelerated training pipeline capable of distributing workloads across multiple graphics processing units through multi-process distribution. It supports the generation of descriptions for both static image files and real-time video streams. The framework includes capabilities for enco

    Neuraltalk2 is an open-source multimodal vision-language model for generating captions from images, squarely fitting the category but limited to captioning rather than broader visual understanding like full question answering or chat.

    Jupyter NotebookImage CaptioningImage Captioning Training
    在 GitHub 上查看↗5,588
  • thudm/cogvlm2THUDM 的头像

    THUDM/CogVLM2

    2,437在 GitHub 上查看↗

    中文版README

    CogVLM2 is an open-source multimodal vision-language model from THUDM that handles image understanding, visual question answering, and image captioning with fine-tuning and zero-shot inference, directly matching the request for a GPT-4 Vision-like model.

    PythonFine-tuned ModelsMultimodal Agents
    在 GitHub 上查看↗2,437
  • google-deepmind/gemmagoogle-deepmind 的头像

    google-deepmind/gemma

    5,475在 GitHub 上查看↗

    Gemma is a family of open-weights large language models based on a decoder-only transformer architecture. These models are designed for text generation and multi-modal conversations, capable of processing and generating responses based on both textual and visual input sequences. The project provides a fine-tunable AI model that supports weight adjustment and low-rank adaptation to specialize performance for particular tasks. It includes support for quantized weights to reduce memory usage and increase inference speed on limited hardware. The capability surface covers multi-modal AI integrati

    Gemma is an open-weights multimodal language model family that processes both text and image inputs, supports fine-tuning and quantized deployment, and is pre-trained for visual tasks—directly matching the need for an open-source vision-language model.

    PythonModel Fine-TuningOpen-Weights Models
    在 GitHub 上查看↗5,475
  • qwenlm/qwen-vlQwenLM 的头像

    QwenLM/Qwen-VL

    6,535在 GitHub 上查看↗

    Qwen-VL is a multimodal large language model that processes both text and images, enabling tasks like image captioning and visual question answering with open-source weights and support for fine-tuning, directly matching the request for an open-source vision-language model.

    PythonVision-Language Models
    在 GitHub 上查看↗6,535
  • deepseek-ai/deepseek-vldeepseek-ai 的头像

    deepseek-ai/DeepSeek-VL

    4,134在 GitHub 上查看↗

    DeepSeek-VL is a multimodal large language model and image-to-text reasoning engine. It functions as a vision-language model and visual question answering system that integrates visual perception with linguistic reasoning to understand and describe images. The project enables multimodal image understanding and document image analysis, specifically processing screenshots of web pages and technical diagrams. It provides capabilities for visual conversational AI, allowing users to interact with visual data to extract insights and perform complex reasoning across different types of visual informa

    DeepSeek-VL is an open-source vision-language model purpose-built for real-world image understanding, covering tasks like captioning and visual question answering, which directly matches your search for a multimodal AI model similar to GPT-4 Vision.

    PythonVision-Language ModelsVision-Language ModelsVisual Question Answering
    在 GitHub 上查看↗4,134
  • internlm/internlm-xcomposerInternLM 的头像

    InternLM/InternLM-XComposer

    2,924在 GitHub 上查看↗

    InternLM-XComposer-2.5

    InternLM-XComposer-2.5 is an open-source vision-language model built to handle multimodal input (text images) and perform tasks like image understanding and reasoning, which directly matches the query for a GPT-4 Vision–like model, though the sparse description leaves some specifics on features like fine-tuning and inference servers unconfirmed.

    PythonVision Language Models
    在 GitHub 上查看↗2,924
  • microsoft/llava-medmicrosoft 的头像

    microsoft/LLaVA-Med

    2,214在 GitHub 上查看↗

    Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.

    LLaVA-Med is a multimodal vision-language model fine-tuned for biomedicine, capable of understanding and analyzing medical images alongside text, which aligns with the goal of an image-understanding AI, though its domain specialization makes it a narrower fit than a general-purpose model.

    PythonMedical Multimodal ModelsMedical Vision Language ModelsPre-training Datasets
    在 GitHub 上查看↗2,214
一览前 10 名对比
仓库Star 数语言许可证最后推送
salesforce/lavis11.2KJupyter NotebookBSD-3-Clause2026年6月2日
thudm/glm-47.1KPythonApache-2.02025年7月4日
vision-cair/minigpt-425.7KPythonBSD-3-Clause2024年9月2日
salesforce/blip5.7KJupyter Notebookbsd-3-clause2024年8月5日
vikhyat/moondream9.8KPythonApache-2.02026年4月20日
clovaai/donut6.8KPythonmit2024年7月11日
opengvlab/internvl10.1KPythonMIT2025年9月22日
haotian-liu/llava24.5KPythonapache-2.02024年8月12日
mlfoundations/open_clip13.9KPythonNOASSERTION2026年6月22日
karpathy/neuraltalk25.6KJupyter Notebook—2017年11月7日

Related searches

  • an open source model for image generation
  • 通用图像分割模型
  • AI 图像描述模型
  • 开源 LLM 交互界面
  • 开源 LLM 托管平台
  • 开源 Copilot 替代方案
  • 开源 LLM 应用开发框架
  • 语言模型对比评测基准