For 开源视觉语言模型, the strongest matches are salesforce/lavis (LAVIS is a comprehensive framework and library of pre-trained), thudm/glm-4 (GLM-4 is an open-weights multimodal language model that processes) and vision-cair/minigpt-4 (MiniGPT-4 is an open-source multimodal vision-language model that processes). salesforce/blip and vikhyat/moondream round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
探索专为高级图像理解和多模态视觉推理任务而设计的开源模型与框架。
LAVIS is a multimodal large language model framework and vision-language model library. It provides tools for training and evaluating models that integrate visual, textual, and audio data, serving as a cross-modal feature extractor and a zero-shot visual reasoning engine. The framework distinguishes itself by using frozen-backbone integration, where pretrained encoders remain non-trainable while lightweight adapter layers are updated. It employs cross-modal feature alignment to map different representations into a shared embedding space and utilizes a modular model wrapper to swap vision and
LAVIS is a comprehensive framework and library of pre-trained vision-language models (like BLIP and ALBEF) that directly supports image captioning, visual question answering, and zero-shot multimodal reasoning, making it a flagship open-source toolbox for your needs.
GLM-4 is an open weights large language model designed as a multimodal chat system. It functions as a reasoning-focused and multilingual model capable of processing and generating responses across text and visual data types. The model is distinguished by its function-calling capabilities, allowing it to interface with external tools and APIs to execute tasks and retrieve real-time information. It is optimized for complex logical reasoning, mathematical problem solving, and deep research involving long-form content generation. Broad capabilities include multilingual text generation, the creat
GLM-4 is an open-weights multimodal language model that processes both text and images, supports fine-tuning and function calling, and is built for reasoning and visual understanding, directly matching the need for an open-source GPT-4 Vision alternative.
MiniGPT-4 is a multimodal AI framework and large language model that integrates vision encoders with language models to process and reason about combined image and text inputs. It functions as a vision-language model capable of image-based conversational AI, visual question answering, and multimodal logical reasoning. The project utilizes a pretrained vision-language integration strategy that connects a vision encoder to a language model via a linear projection layer. This approach employs frozen-backbone training to align visual representations with linguistic tokens while keeping the primar
MiniGPT-4 is an open-source multimodal vision-language model that processes text and images together, supporting visual question answering, image captioning, and conversational reasoning, with pre-trained weights and fine-tuning capability — exactly the kind of model this search is after.
BLIP is a vision-language model framework that combines contrastive, matching, and language modeling objectives to align images with text. Built on a multimodal encoder-decoder architecture, it supports distributed data-parallel training with cosine learning rate scheduling and sliding-window metric tracking for training stability. The framework provides capabilities for image captioning, visual question answering, and cross-modal retrieval, scoring semantic alignment between images and text through learned embeddings. It includes toolkits for fine-tuning pre-trained models on custom datasets
BLIP is a multimodal vision-language model framework that handles image captioning, visual question answering, and cross-modal retrieval, with open-source weights, fine-tuning support, and pre-training on large datasets — exactly the kind of image-understanding AI you’re looking for.
Moondream is a small-scale vision language model designed to reason across images to generate captions and answer natural language questions. It functions as an edge-optimized system capable of performing visual question answering, image captioning, and object detection. The project distinguishes itself through a lightweight architecture designed for local inference on embedded devices, workstations, and air-gapped hardware. It supports the execution of models on local GPUs and Apple Silicon to ensure data privacy and low latency. The system's capabilities include identifying precise object
Moondream is an open-source vision-language model that generates captions and answers questions about images, supporting local inference and fine-tuning — exactly the multimodal image understanding tool you are looking for.
Donut is an OCR-free document transformer and end-to-end document parser. It functions as a neural network that converts unstructured document images directly into structured data or text without the use of an external optical character recognition engine. The project includes a synthetic document generator to create artificial images and ground-truth labels for training. It employs a transformer model to perform visual question answering and document image classification based on visual layout and text. The system covers several document understanding capabilities, including structured info
Donut is a multimodal vision-language model specialized in document understanding — it supports visual question answering and captioning on document images, but its domain focus means it may not handle general images like GPT-4 Vision.
InternVL is a vision-language model framework that fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning. It provides a system for multimodal inference and dialogue, enabling the processing of images and text to answer questions or generate descriptions. The project is distinguished by its high-resolution image processing, which uses dynamic tiling to maintain detail for images up to 4K resolution, and its chain-of-thought visual reasoning for solving complex mathematical and spatial problems. It also supports temporal frame sampling
InternVL is an open-source vision-language model that fuses a visual encoder with a large language model for multimodal reasoning, covering image captioning, visual question answering, and high-resolution image processing, directly matching the search for an open-source alternative to GPT-4 Vision.
LLaVA is a multimodal large language model architecture designed to process and interpret both image and text inputs to generate natural language responses. It functions as a research-oriented platform for visual instruction tuning, providing a framework to align language models with human intent through training on diverse datasets of paired images and text queries. The system distinguishes itself through a specialized vision-language training pipeline that connects visual data to language models using projection layers and instruction-based fine-tuning. It supports distributed inference by
LLaVA is an open-source multimodal vision-language model that processes both image and text inputs to generate natural language responses, supporting image captioning, visual question answering, fine-tuning, and zero-shot inference — squarely matching the need for a GPT-4 Vision alternative.
Open CLIP is an open source framework for training and deploying Contrastive Language-Image Pre-training models. It serves as a vision-language training framework and multimodal embedding engine that maps images and text into a shared vector space for similarity searches and zero-shot classification. The project provides a toolkit for distributed training of contrastive models and includes an image-to-text generative model for producing natural language descriptions. It supports custom text encoder integration and utilizes teacher-student model distillation to transfer knowledge from large pr
Open CLIP is a framework for training and deploying contrastive vision-language models like CLIP, a foundational open-source multimodal model that maps images and text into a shared space for zero-shot classification and captioning — it covers image understanding, fine-tuning, and pre-trained weights, though it is more of an embedding engine than a generative VQA system like GPT-4V.
Neuraltalk2 is a deep learning vision system designed for automatic image captioning. Built with PyTorch, it utilizes a hybrid architecture that combines a convolutional neural network encoder with a recurrent neural network decoder to generate textual descriptions from visual input. The project features a GPU-accelerated training pipeline capable of distributing workloads across multiple graphics processing units through multi-process distribution. It supports the generation of descriptions for both static image files and real-time video streams. The framework includes capabilities for enco
Neuraltalk2 is an open-source multimodal vision-language model for generating captions from images, squarely fitting the category but limited to captioning rather than broader visual understanding like full question answering or chat.
中文版README
CogVLM2 is an open-source multimodal vision-language model from THUDM that handles image understanding, visual question answering, and image captioning with fine-tuning and zero-shot inference, directly matching the request for a GPT-4 Vision-like model.
Gemma is a family of open-weights large language models based on a decoder-only transformer architecture. These models are designed for text generation and multi-modal conversations, capable of processing and generating responses based on both textual and visual input sequences. The project provides a fine-tunable AI model that supports weight adjustment and low-rank adaptation to specialize performance for particular tasks. It includes support for quantized weights to reduce memory usage and increase inference speed on limited hardware. The capability surface covers multi-modal AI integrati
Gemma is an open-weights multimodal language model family that processes both text and image inputs, supports fine-tuning and quantized deployment, and is pre-trained for visual tasks—directly matching the need for an open-source vision-language model.
Qwen-VL is a multimodal large language model that processes both text and images, enabling tasks like image captioning and visual question answering with open-source weights and support for fine-tuning, directly matching the request for an open-source vision-language model.
DeepSeek-VL is a multimodal large language model and image-to-text reasoning engine. It functions as a vision-language model and visual question answering system that integrates visual perception with linguistic reasoning to understand and describe images. The project enables multimodal image understanding and document image analysis, specifically processing screenshots of web pages and technical diagrams. It provides capabilities for visual conversational AI, allowing users to interact with visual data to extract insights and perform complex reasoning across different types of visual informa
DeepSeek-VL is an open-source vision-language model purpose-built for real-world image understanding, covering tasks like captioning and visual question answering, which directly matches your search for a multimodal AI model similar to GPT-4 Vision.
InternLM-XComposer-2.5
InternLM-XComposer-2.5 is an open-source vision-language model built to handle multimodal input (text images) and perform tasks like image understanding and reasoning, which directly matches the query for a GPT-4 Vision–like model, though the sparse description leaves some specifics on features like fine-tuning and inference servers unconfirmed.
Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.
LLaVA-Med is a multimodal vision-language model fine-tuned for biomedicine, capable of understanding and analyzing medical images alongside text, which aligns with the goal of an image-understanding AI, though its domain specialization makes it a narrower fit than a general-purpose model.
| 仓库 | Star 数 | 语言 | 许可证 | 最后推送 |
|---|---|---|---|---|
| salesforce/lavis | 11.2K | Jupyter Notebook | BSD-3-Clause | |
| thudm/glm-4 | 7.1K | Python | Apache-2.0 | |
| vision-cair/minigpt-4 | 25.7K | Python | BSD-3-Clause | |
| salesforce/blip | 5.7K | Jupyter Notebook | bsd-3-clause | |
| vikhyat/moondream | 9.8K | Python | Apache-2.0 | |
| clovaai/donut | 6.8K | Python | mit | |
| opengvlab/internvl | 10.1K | Python | MIT | |
| haotian-liu/llava | 24.5K | Python | apache-2.0 | |
| mlfoundations/open_clip | 13.9K | Python | NOASSERTION | |
| karpathy/neuraltalk2 | 5.6K | Jupyter Notebook | — |