awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com

Vision language models

Ranking updated Sep 7, 2026

For vision language models, the first results are haotian-liu/llava (This repository provides a foundational multimodal large language model architecture with a dedicated vision encoder and instruction-tuning support, squarely matching the requirement for vision-language processing and reasoning), deepseek-ai/deepseek-vl2 (DeepSeek-VL2 is a multimodal vision-language model equipped with visual reasoning, object grounding, document understanding, and efficiency optimizations, making it a comprehensive fit for this search) and qwenlm/qwen2-vl (Qwen2-VL is a multimodal large language model designed for visual reasoning across images and video, directly providing object detection, video understanding, and fine-tuning capabilities for vision-language tasks). deepseek-ai/deepseek-vl and openbmb/minicpm-v round out the shortlist. Compare the match explanations and check the project documentation against your requirements.

Compare the top open-source vision language models for your AI stack. Hand-picked repositories ranked by stars, activity, and capabilities. Find the best fit.

Vision language models

Find the best repos with AI.We'll search the best matching repositories with AI.
  • haotian-liu/llavahaotian-liu avatar

    haotian-liu/LLaVA

    24,465View on GitHub↗

    LLaVA is a multimodal large language model architecture designed to process and interpret both image and text inputs to generate natural language responses. It functions as a research-oriented platform for visual instruction tuning, providing a framework to align language models with human intent through training on diverse datasets of paired images and text queries. The system distinguishes itself through a specialized vision-language training pipeline that connects visual data to language models using projection layers and instruction-based fine-tuning. It supports distributed inference by

    This repository provides a foundational multimodal large language model architecture with a dedicated vision encoder and instruction-tuning support, squarely matching the requirement for vision-language processing and reasoning.

    PythonModel Fine-TuningVision EncodersMultimodal Large Language Models
    View on GitHub↗24,465
  • deepseek-ai/deepseek-vl2deepseek-ai avatar

    deepseek-ai/DeepSeek-VL2

    5,302View on GitHub↗

    DeepSeek-VL2 is a multimodal large language model and vision-language system designed to analyze visual scenes and generate descriptive text. It functions as a visual question answering and visual grounding model, capable of extracting information from documents and locating specific objects or regions within images based on textual descriptions. The project utilizes a mixture-of-experts architecture to process combined image and text inputs. It is optimized for inference through incremental prefilling, which reduces the GPU memory requirements on hardware. The model covers multimodal data a

    DeepSeek-VL2 is a multimodal vision-language model equipped with visual reasoning, object grounding, document understanding, and efficiency optimizations, making it a comprehensive fit for this search.

    PythonMultimodal Large Language ModelsVision-Language Grounding ModelsVisual Object Grounding
    View on GitHub↗5,302
  • qwenlm/qwen2-vlQwenLM avatar

    QwenLM/Qwen2-VL

    19,404View on GitHub↗

    Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw

    Qwen2-VL is a multimodal large language model designed for visual reasoning across images and video, directly providing object detection, video understanding, and fine-tuning capabilities for vision-language tasks.

    Jupyter NotebookObject DetectionMultimodal Large Language ModelsVideo Understanding Models
    View on GitHub↗19,404
  • deepseek-ai/deepseek-vldeepseek-ai avatar

    deepseek-ai/DeepSeek-VL

    4,134View on GitHub↗

    DeepSeek-VL is a multimodal large language model and image-to-text reasoning engine. It functions as a vision-language model and visual question answering system that integrates visual perception with linguistic reasoning to understand and describe images. The project enables multimodal image understanding and document image analysis, specifically processing screenshots of web pages and technical diagrams. It provides capabilities for visual conversational AI, allowing users to interact with visual data to extract insights and perform complex reasoning across different types of visual informa

    DeepSeek-VL is an open-source vision-language model equipped with a vision encoder and multimodal reasoning capabilities, though it lacks dedicated video understanding and object grounding features.

    PythonMultimodal Visual ReasoningVision Transformer EncodersMultimodal
    View on GitHub↗4,134
  • openbmb/minicpm-vOpenBMB avatar

    OpenBMB/MiniCPM-V

    25,653View on GitHub↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    MiniCPM-V is a compact vision-language model supporting multimodal reasoning, video understanding, efficient quantization, and fine-tuning for edge devices, directly matching the search for open-source visual-text models.

    PythonModel QuantizationMultimodal Large Language ModelsVideo Understanding Models
    View on GitHub↗25,653
  • llava-vl/llava-nextLLaVA-VL avatar

    LLaVA-VL/LLaVA-NeXT

    4,695View on GitHub↗

    LLaVA-NeXT is a multimodal large language model framework and training toolkit designed to process interleaved images and video sequences to generate text. It functions as a visual language model that combines vision encoders with language models to perform complex reasoning, question answering, and video understanding. The system is capable of analyzing high-resolution images and temporal video frames to describe events, summarize actions, and reason across multiple visual inputs. It supports the interpretation of documents and charts, spatial environment analysis, and the generation of desc

    LLaVA-NeXT is a multimodal vision-language framework designed for complex reasoning across interleaved images and video sequences, perfectly matching your search for vision-language models with fine-tuning and video capabilities.

    PythonMultimodal Visual ReasoningVideo UnderstandingVideo Understanding Models
    View on GitHub↗4,695
  • apple/ml-ferretapple avatar

    apple/ml-ferret

    8,680View on GitHub↗

    ml-ferret is a multimodal large language model framework and visual reasoning engine designed to reason about images and user interfaces. It functions as a UI grounding model and referring expression comprehension tool that maps natural language descriptions to precise pixel coordinates. The system focuses on high-resolution image analysis to identify and locate specific interface components. It employs multi-resolution image processing and region-aware visual encoding to preserve detail across different aspect ratios, enabling the model to analyze spatial relationships and functional layouts

    This repository provides a multimodal large language model framework tailored for visual reasoning and grounding, covering several required features like multimodal reasoning and vision encoding, though it lacks direct support for video understanding.

    PythonMultimodal Large Language ModelsVision-Language Grounding ModelsVisual Grounding
    View on GitHub↗8,680
  • qwenlm/qwen-vlQwenLM avatar

    QwenLM/Qwen-VL

    6,535View on GitHub↗

    This repository provides a foundational vision-language framework with multimodal reasoning, object grounding, fine-tuning support, and efficient quantization capabilities to process both visual and textual inputs.

    PythonModel QuantizationMultimodal Large Language Models
    View on GitHub↗6,535
  • zai-org/cogvlmzai-org avatar

    zai-org/CogVLM

    6,742View on GitHub↗

    CogVLM is a multimodal large language model designed for visual reasoning and multi-turn dialogue. It functions as a visual grounding model and a quantized vision model, combining text and image processing to perform complex understanding and maintain context across visual inputs. The project includes capabilities as a GUI automation agent, allowing it to analyze application screenshots, plan operational steps, and return precise screen coordinates for interface interaction. It further supports visual grounding by generating bounding box coordinates to map text descriptions to specific spatia

    CogVLM is a vision-language model that provides multimodal reasoning, visual grounding, and object detection capabilities for image and text inputs, though it lacks native video understanding features.

    PythonMultimodal Large Language ModelsVision-Language Grounding ModelsVisual Grounding
    View on GitHub↗6,742
  • opengvlab/internvlOpenGVLab avatar

    OpenGVLab/InternVL

    10,061View on GitHub↗

    InternVL is a vision-language model framework that fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning. It provides a system for multimodal inference and dialogue, enabling the processing of images and text to answer questions or generate descriptions. The project is distinguished by its high-resolution image processing, which uses dynamic tiling to maintain detail for images up to 4K resolution, and its chain-of-thought visual reasoning for solving complex mathematical and spatial problems. It also supports temporal frame sampling

    InternVL is a comprehensive vision-language model framework that combines a robust visual encoder with a large language model to support high-resolution image processing, video understanding, and advanced multimodal reasoning.

    PythonModel Fine-TuningModel QuantizationVideo Understanding
    View on GitHub↗10,061
  • evolvinglmms-lab/otterEvolvingLMMs-Lab avatar

    EvolvingLMMs-Lab/Otter

    3,331View on GitHub↗

    Otter is a framework and toolkit for the pretraining, fine-tuning, and evaluation of vision-language models. It provides a pipeline for training large language models to process high-resolution images and video frames, integrating visual encoders with textual token spaces. The system is designed for multi-visual input processing, allowing models to interpret multiple images or video sequences within a single prompt. It supports multi-round conversation management to maintain context across interactions for detailed scene comprehension and visual reasoning. The framework covers a full develop

    Otter is an open-source vision-language framework designed for multimodal pretraining, fine-tuning, and reasoning across images, videos, and text inputs.

    PythonMultimodal
    View on GitHub↗3,331
  • vision-cair/minigpt-4Vision-CAIR avatar

    Vision-CAIR/MiniGPT-4

    25,679View on GitHub↗

    MiniGPT-4 is a multimodal AI framework and large language model that integrates vision encoders with language models to process and reason about combined image and text inputs. It functions as a vision-language model capable of image-based conversational AI, visual question answering, and multimodal logical reasoning. The project utilizes a pretrained vision-language integration strategy that connects a vision encoder to a language model via a linear projection layer. This approach employs frozen-backbone training to align visual representations with linguistic tokens while keeping the primar

    MiniGPT-4 is a vision-language framework that integrates a pretrained vision encoder with a large language model to handle image-based conversational reasoning, though it lacks dedicated native support for object detection and video understanding.

    PythonMultimodal Reasoning TasksMultimodal Large Language Models
    View on GitHub↗25,679
  • paddlepaddle/larkPaddlePaddle avatar

    PaddlePaddle/LARK

    7,717View on GitHub↗

    LARK is a development toolkit for training, fine-tuning, and deploying large language models and multimodal models based on PaddlePaddle. It functions as a comprehensive framework that includes an LLM training orchestrator, an inference server, and a multimodal model framework for processing text, image, and video inputs. The project features a retrieval-augmented generation system for building conversational applications that integrate web search and private knowledge bases. It provides specific capabilities for multimodal reasoning and complex logic, enabling the extraction of structured da

    This development toolkit supports training, fine-tuning, and deploying multimodal models for text, image, and video inputs, fitting the vision-language model category despite lacking explicit mention of object detection and grounding features.

    PythonModel Fine-TuningModel QuantizationMultimodal Reasoning Tasks
    View on GitHub↗7,717
  • meta-llama/llama-modelsmeta-llama avatar

    meta-llama/llama-models

    7,643View on GitHub↗

    This project provides a foundational framework and reference implementation for executing causal language modeling and multimodal reasoning on local systems. It includes a set of core components for managing model assets, a fine-tuning framework, and structural definitions required to instantiate transformer-based architectures. The system is distinguished by its ability to process combined text and image inputs through multimodal transformer models for visual reasoning and document analysis. It also supports the deployment of quantized models, reducing memory footprints through low-precision

    This repository provides the reference framework and implementation for the Llama family, including the multimodal models required for processing both text and image inputs with fine-tuning and quantization support.

    PythonModel QuantizationMultimodal Visual ReasoningVision-Language Grounding Models
    View on GitHub↗7,643
  • google-research/big_visiongoogle-research avatar

    google-research/big_vision

    3,363View on GitHub↗

    This project is a research framework and toolkit designed for training large-scale vision transformers and multimodal language models. It provides a comprehensive suite for vision-language pretraining, enabling the development of models that map images and text into shared latent spaces. The framework is distinguished by its capabilities in high-fidelity image generation and multimodal research, utilizing normalizing flows and variational autoencoders to produce images from text prompts or class labels. It supports the development of both generative and contrastive models, allowing for a wide

    This research framework and toolkit provides the infrastructure for training large-scale vision transformers and multimodal language models that map images and text into shared spaces, matching the search requirements.

    Jupyter NotebookDistributed Training ShardingLarge Scale TrainingConditional Image Generation
    View on GitHub↗3,363
  • microsoft/omniparsermicrosoft avatar

    microsoft/OmniParser

    24,377View on GitHub↗

    OmniParser is a multimodal interaction engine designed to function as a desktop automation agent. It interprets visual screen information to execute complex, multi-step tasks across operating system environments by bridging visual interface perception with language models. Through a continuous cycle of observation and command execution, the system grounds high-level natural language instructions into precise, coordinate-based actions. The project distinguishes itself by utilizing vision-based parsing to interact with software interfaces without requiring access to underlying application progr

    OmniParser provides a vision-language grounding engine capable of visual interface perception and coordinate-based action, making it a strong tool for multimodal UI reasoning even though it focuses specifically on desktop automation rather than general-purpose video and image understanding.

    Jupyter NotebookVision-Language Grounding Models
    View on GitHub↗24,377
  • thudm/cogvlmTHUDM avatar

    THUDM/CogVLM

    6,742View on GitHub↗

    CogVLM is a multimodal large language model designed to integrate visual and textual data for reasoning about images and generating natural language. It functions as a visual question answering system that analyzes image content to provide detailed descriptions or answer specific questions. The project includes a visual grounding model capable of mapping text descriptions to precise bounding box coordinates within an image. It also features a vision-based automation agent that analyzes screen captures to generate execution plans and interaction coordinates for software interfaces. The system

    CogVLM is a multimodal vision-language model supporting image reasoning, visual grounding, and fine-tuning, though it focuses primarily on images rather than full video understanding.

    PythonMultimodal Large Language Models
    View on GitHub↗6,742
  • thudm/visualglm-6bTHUDM avatar

    THUDM/VisualGLM-6B

    4,157View on GitHub↗

    VisualGLM-6B is a bilingual multimodal large language model and vision-language model designed for conversational tasks and visual understanding. It functions as a bilingual AI model capable of processing and generating responses in both Chinese and English. The system is a quantized large language model supporting 4-bit and 8-bit precision to reduce memory usage and hardware requirements during local deployment. It is also a parameter-efficient fine-tuning model, allowing for weight adjustments to adapt the system to specific downstream tasks without full retraining. The project covers mult

    VisualGLM-6B is an open-source multimodal large language model equipped with a vision encoder for bilingual image-text conversation, supporting both quantization and fine-tuning.

    PythonMultimodal Large Language Models
    View on GitHub↗4,157
  • apple/ml-fastvlmapple avatar

    apple/ml-fastvlm

    7,375View on GitHub↗

    This project is a vision language model framework and vision-to-text pipeline designed for deploying and optimizing models that process both images and text. It provides an on-device inference engine and a vision language model framework to run quantized models locally on mobile and desktop hardware accelerators. The framework features a model quantization toolkit to reduce weight precision for lower memory footprints and increased execution speed on specialized silicon. It also includes an efficient vision encoder utilizing a hybrid encoding system to compress image tokens, which reduces pro

    This repository provides an on-device vision-language model framework and quantization toolkit designed for efficient image and text processing on hardware accelerators, though it lacks explicit grounding and extensive video understanding features.

    PythonModel Quantization
    View on GitHub↗7,375
  • vectorspacelab/omnigen2VectorSpaceLab avatar

    VectorSpaceLab/OmniGen2

    4,093View on GitHub↗

    OmniGen2 is a unified image generation model and multimodal large language model designed to handle text-to-image generation, image-to-image tasks, and image editing within a single framework. It functions as a causal language model visual engine capable of generating and editing images based on combined text and visual inputs. The system features in-context visual composition and subject-driven generation, allowing it to extract subjects from reference images and place them into new scenes. It also supports instruction-based image editing, where specific objects or styles are modified via na

    OmniGen2 is a unified vision-language model that handles multimodal reasoning and visual editing through a causal language model framework, though its primary focus is on generative tasks rather than general video understanding or object grounding.

    Jupyter NotebookMultimodal Visual Reasoning
    View on GitHub↗4,093
  • clovaai/donutclovaai avatar

    clovaai/donut

    6,789View on GitHub↗

    Donut is an OCR-free document transformer and end-to-end document parser. It functions as a neural network that converts unstructured document images directly into structured data or text without the use of an external optical character recognition engine. The project includes a synthetic document generator to create artificial images and ground-truth labels for training. It employs a transformer model to perform visual question answering and document image classification based on visual layout and text. The system covers several document understanding capabilities, including structured info

    Donut is a vision-language transformer designed for OCR-free document understanding and multimodal document parsing, which fits the category well despite being specialized for documents rather than general video and image reasoning.

    PythonModel Fine-Tuning
    View on GitHub↗6,789
  • microsoft/unilmmicrosoft avatar

    microsoft/unilm

    22,030View on GitHub↗

    This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec

    This project provides a comprehensive foundation and toolkit for transformer-based multimodal models that handle visual and textual inputs, though it encompasses a broader family of architectures rather than a single turnkey vision-language model.

    PythonModel Fine-TuningMultimodal Large Language ModelsVision-Language Grounding Models
    View on GitHub↗22,030
  • huggingface/transformershuggingface avatar

    huggingface/transformers

    161,630View on GitHub↗

    Transformers is a comprehensive library for machine learning that provides a unified interface for training, fine-tuning, and deploying transformer-based models. It supports a wide range of tasks, including text classification, language modeling, question answering, and sequence-to-sequence translation, while offering specialized architectures for both text and vision processing. The framework includes tools for managing the entire model lifecycle, from data preprocessing and tokenization to distributed training and inference. The library features extensive support for model optimization and

    Transformers is a foundational machine learning framework that provides the underlying model architectures and pipelines to load, train, and run vision-language models, though it is a general-purpose library rather than a pre-packaged multimodal model itself.

    PythonModel Quantization
    View on GitHub↗161,630
  • sgl-project/sglangsgl-project avatar

    sgl-project/sglang

    29,079View on GitHub↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Sglang is a high-performance inference engine and serving system for large language and multimodal models, making it the right kind of tool for running vision-language models in production even though it focuses on serving and orchestration rather than training or base model architectures.

    PythonModel QuantizationModel Quantization
    View on GitHub↗29,079
  • autogluon/autogluonautogluon avatar

    autogluon/autogluon

    9,997View on GitHub↗

    AutoGluon is an automated machine learning framework and multimodal library designed to automate the end-to-end pipeline from data preprocessing to high-accuracy model training and validation. It functions as an automated model trainer for tabular, image, text, and time series data, as well as a tool for time series forecasting and foundation model finetuning. The project is distinguished by its ability to jointly process and fuse different data types, allowing for the construction of multimodal neural networks that integrate images, text, and structured tables. It supports zero-shot inferenc

    AutoGluon is an automated machine learning framework that builds multimodal models combining image, text, and tabular data, aligning well with the search for vision-language capabilities though it functions as an AutoML pipeline rather than a standalone foundation model.

    PythonModel Fine-TuningObject Detection
    View on GitHub↗9,997
  • hiyouga/llama-factoryhiyouga avatar

    hiyouga/LLaMA-Factory

    72,241View on GitHub↗

    LLaMA-Factory is a comprehensive suite for dataset preparation, model fine-tuning, memory optimization, and standardized API deployment. It provides a unified platform for the supervised and reward-based fine-tuning of large language models and vision-language models. The framework includes a specialized toolkit for training vision-language models and a model serving interface that deploys trained models through high-performance APIs. It utilizes precision tuning and quantization techniques to reduce the hardware requirements and memory footprint of large models. The system covers data pipel

    LLaMA-Factory is a fine-tuning and training framework for large language and vision-language models, though it serves as a training toolkit rather than providing pre-trained multimodal models out of the box.

    PythonLarge Language Model Fine-Tuning FrameworksDataset Preparation ToolsDistributed Training
    View on GitHub↗72,241
  • jingyaogong/minimindjingyaogong avatar

    jingyaogong/minimind

    51,834View on GitHub↗

    This project is a comprehensive framework for the entire lifecycle of transformer-based language models, supporting everything from foundational pretraining to specialized deployment. It provides a modular toolkit for defining neural network architectures, managing data preparation pipelines, and executing training routines across various scales. The framework is designed to handle the full model development process, including supervised fine-tuning, behavioral alignment, and the integration of agentic capabilities. What distinguishes this framework is its focus on efficient training and adva

    This repository provides a modular framework for training and fine-tuning transformer-based language models, though its primary focus is on text rather than fully integrated vision-language processing.

    PythonModel Training ToolkitsAgentic FrameworksAgentic Training Frameworks
    View on GitHub↗51,834
  • ml-gsai/lladaML-GSAI avatar

    ML-GSAI/LLaDA

    3,580View on GitHub↗

    LLaDA is a masked diffusion language model and conditional text generator. It generates text by iteratively refining masked tokens through a diffusion process rather than predicting the next token in a sequence. The project functions as a vision-language diffusion model, converting visual inputs into text responses. It also serves as a preference optimization framework that uses log-likelihood estimation and evidence lower bounds to tune model responses. The system supports multi-round conversational AI and text sequence evaluation. It integrates vision-language embedding for cross-modal con

    LLaDA is a vision-language diffusion model that integrates cross-attention fusion for visual-to-text generation and multi-turn interaction, though it lacks some specific grounding and efficiency features.

    PythonMasked Text DiffusionMultimodal Diffusion ModelsCross-Attention Conditioning
    View on GitHub↗3,580
  • openai/clipopenai avatar

    openai/CLIP

    33,779View on GitHub↗

    CLIP is a neural network architecture designed to map visual and textual data into a shared latent vector space. By utilizing transformer-based feature extraction and multi-modal tokenization, the system aligns images and natural language strings, enabling cross-modal similarity analysis and semantic classification. The project functions as a zero-shot classification engine, identifying image content by calculating the cosine similarity between visual features and arbitrary text labels without requiring task-specific retraining. Beyond inference, it serves as a research toolkit for evaluating

    This repository provides a foundational multimodal neural network that aligns images and text into a shared latent space for zero-shot classification and cross-modal retrieval, though it lacks direct support for complex video understanding or object grounding.

    Jupyter NotebookContrastive Learning ModelsZero-Shot Inference EnginesComputer Vision Evaluation Tools
    View on GitHub↗33,779
  • bytedance-seed/bagelByteDance-Seed avatar

    ByteDance-Seed/Bagel

    5,681View on GitHub↗

    This repository provides vision-language models and unified generation tools that handle multimodal tasks, though it lacks explicit out-of-the-box mentions of object detection and grounding.

    PythonImage EditingAutoregressive Visual Token PredictorsCausal Masked Modeling Objectives
    View on GitHub↗5,681
  • ailab-cvc/yolo-worldAILab-CVC avatar

    AILab-CVC/YOLO-World

    6,425View on GitHub↗

    YOLO-World is a vision-language framework and open-vocabulary object detection model. It identifies objects in images and video based on free-form text prompts without requiring predefined category labels. The system enables the identification of arbitrary objects by fusing image features with text embeddings. It includes a specialized tool for automated image labeling, which generates bounding box annotations for custom datasets using text-based prompts. The project provides a deployment pipeline for converting models into quantized ONNX and TFLite formats, supporting real-time inference on

    YOLO-World is an open-vocabulary vision-language model for object detection and grounding using text prompts, making it a relevant fit despite focusing more on detection than general multimodal reasoning.

    PythonOpen-Vocabulary DetectionOpen-Vocabulary Object DetectionVision-Language Cross-Attention Fusions
    View on GitHub↗6,425
  • salesforce/blipsalesforce avatar

    salesforce/BLIP

    5,676View on GitHub↗

    BLIP is a vision-language model framework that combines contrastive, matching, and language modeling objectives to align images with text. Built on a multimodal encoder-decoder architecture, it supports distributed data-parallel training with cosine learning rate scheduling and sliding-window metric tracking for training stability. The framework provides capabilities for image captioning, visual question answering, and cross-modal retrieval, scoring semantic alignment between images and text through learned embeddings. It includes toolkits for fine-tuning pre-trained models on custom datasets

    This project is a vision-language model framework providing multimodal alignment, image captioning, visual reasoning, and fine-tuning support, though it lacks dedicated object detection and deep video understanding features.

    Jupyter NotebookEncoder-Decoder ArchitecturesTraining FrameworksCross-Modal Similarity Scoring
    View on GitHub↗5,676
  • airaria/visual-chinese-llama-alpacaairaria avatar

    airaria/Visual-Chinese-LLaMA-Alpaca

    460View on GitHub↗

    多模态中文LLaMA&Alpaca大语言模型(VisualCLA)

    This repository provides a multimodal Chinese vision-language model built on LLaMA and Alpaca with fine-tuning support, making it the right kind of tool despite lacking some advanced features like video understanding.

    PythonMultimodal LLM Models
    View on GitHub↗460
  • minimax-ai/minimax-01MiniMax-AI avatar

    MiniMax-AI/MiniMax-01

    3,435View on GitHub↗

    The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention

    This repository provides the MiniMax-VL-01 vision-language model built on linear attention, making it the right kind of tool for multimodal tasks even though specific grounding and quantization features are not explicitly detailed.

    PythonAttention OptimizationModel ArchitecturesMultimodal LLM Models
    View on GitHub↗3,435
  • openbmb/viscpmOpenBMB avatar

    OpenBMB/VisCPM

    1,068View on GitHub↗

    ICLR'24 spotlight Chinese and English Multimodal Large Model Series (Chat and Paint) | 基于CPM基础模型的中英双语多模态大模型系列

    This repository provides bilingual multimodal large models capable of processing both visual and textual inputs, fitting the requested category though it is narrower than the full feature set.

    PythonMultimodal LLM Models
    View on GitHub↗1,068
  • tensorflow/modelstensorflow avatar

    tensorflow/models

    77,663View on GitHub↗

    This repository serves as a centralized collection of state-of-the-art deep learning architectures and reference implementations designed for research and application development. It provides a comprehensive toolkit for computer vision and natural language processing, offering pre-built models and training pipelines for tasks ranging from image classification and object detection to complex sequence modeling. The project distinguishes itself by providing a flexible execution harness that manages the entire training lifecycle, including data ingestion and backpropagation. It supports scalable

    This repository provides a broad collection of deep learning architectures and reference implementations for computer vision and natural language processing, though it lacks a unified multimodal vision-language framework tailored specifically for joint reasoning across image, video, and text inputs.

    PythonComputer Vision ModelsDevelopment and Orchestration ToolsDistributed Parameter Synchronisation
    View on GitHub↗77,663
  • qwenlm/qwen3-vlQwenLM avatar

    QwenLM/Qwen3-VL

    18,329View on GitHub↗

    Qwen3-VL is a multimodal vision-language model designed to process and reason across images, videos, and text. It functions as a computer vision framework capable of identifying objects, extracting structured data from documents, and interpreting spatial elements within visual media. The system operates as an automated user interface interaction agent, interpreting screen data to navigate software and mobile applications. By utilizing a unified transformer architecture, it performs complex visual reasoning to execute user-defined tasks without manual input. Beyond interface navigation, the m

    Qwen3-VL is an open-source vision-language model designed for multimodal reasoning across images, videos, and text, fitting the requested category well despite missing explicit mentions of efficient quantization in the provided description.

    Jupyter NotebookVision-Language ModelsComputer VisionWeb Interaction Agents
    View on GitHub↗18,329
  • osilly/vision-r1Osilly avatar

    Osilly/Vision-R1

    1,475View on GitHub↗

    The official repo for "Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models".

    Vision-R1 is a multimodal large language model designed to incentivize reasoning capabilities across visual and textual inputs, though it lacks a full suite of object detection and quantization features out of the box.

    PythonMultimodal UnderstandingReasoning Models
    View on GitHub↗1,475
  • salesforce/lavissalesforce avatar

    salesforce/LAVIS

    11,236View on GitHub↗

    LAVIS is a multimodal large language model framework and vision-language model library. It provides tools for training and evaluating models that integrate visual, textual, and audio data, serving as a cross-modal feature extractor and a zero-shot visual reasoning engine. The framework distinguishes itself by using frozen-backbone integration, where pretrained encoders remain non-trainable while lightweight adapter layers are updated. It employs cross-modal feature alignment to map different representations into a shared embedding space and utilizes a modular model wrapper to swap vision and

    This repository provides a modular multimodal framework and vision-language library supporting cross-modal reasoning, feature extraction, and model training, though it lacks dedicated object detection and video understanding pipelines.

    Jupyter NotebookMultimodal FrameworksMultimodal ModelsAdapter Layers
    View on GitHub↗11,236
  • wangclnlp/vision-llm-alignmentwangclnlp avatar

    wangclnlp/Vision-LLM-Alignment

    3View on GitHub↗

    This repository contains the code for SFT, RLHF, and DPO, designed for vision-based LLMs, including the LLaVA models and the LLaMA-3.2-vision models.

    This repository provides alignment and fine-tuning tools for vision-language models like LLaVA and LLaMA-3.2-vision, offering the requested fine-tuning support even though it is a training toolkit rather than a base model itself.

    PythonAlignment and RLHF
    View on GitHub↗3
  • yuyq96/r1-visionyuyq96 avatar

    yuyq96/R1-Vision

    48View on GitHub↗

    R1-Vision: Let's first take a look at the image

    This repository provides a vision-language reasoning model designed to interpret images, though it lacks dedicated object detection and fine-tuning support.

    PythonReasoning Models
    View on GitHub↗48
Compare the top 10 at a glance
RepositoryStarsLanguageLicenseLast push
haotian-liu/llava24.5KPythonapache-2.0Aug 12, 2024
deepseek-ai/deepseek-vl25.3KPythonMITFeb 26, 2025
qwenlm/qwen2-vl19.4KJupyter NotebookApache-2.0Jan 30, 2026
deepseek-ai/deepseek-vl4.1KPythonMITApr 24, 2024
openbmb/minicpm-v25.7KPythonApache-2.0Jun 4, 2026
llava-vl/llava-next4.7KPythonApache-2.0Jun 15, 2026
apple/ml-ferret8.7KPythonNOASSERTIONOct 9, 2024
qwenlm/qwen-vl6.5KPythonotherAug 6, 2024
zai-org/cogvlm6.7KPythonApache-2.0May 29, 2024
opengvlab/internvl10.1KPythonMITSep 22, 2025

Related searches

  • an open-source vision-language model
  • Computer vision and multimodal
  • Vision language models
  • a model for AI image captioning
  • a computer vision library for Python
  • a benchmark for comparing language models
  • a comprehensive collection of LLM research papers
  • a computer vision library for object detection