awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

4 repositorios

Awesome GitHub RepositoriesVision Language Model

Explore 4 awesome GitHub repositories matching part of an awesome list · Vision Language Model. Refine with filters or upvote what's useful.

Awesome Vision Language Model GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • haotian-liu/llavaAvatar de haotian-liu

    haotian-liu/LLaVA

    24,465Ver en GitHub↗

    LLaVA is a multimodal large language model architecture designed to process and interpret both image and text inputs to generate natural language responses. It functions as a research-oriented platform for visual instruction tuning, providing a framework to align language models with human intent through training on diverse datasets of paired images and text queries. The system distinguishes itself through a specialized vision-language training pipeline that connects visual data to language models using projection layers and instruction-based fine-tuning. It supports distributed inference by

    Listed in the “Vision Language Model” section of the Ailia Models awesome list.

    Pythonchatbotchatgptfoundation-models
    Ver en GitHub↗24,465
  • qwenlm/qwen2-vlAvatar de QwenLM

    QwenLM/Qwen2-VL

    19,404Ver en GitHub↗

    Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw

    Listed in the “Vision Language Model” section of the Ailia Models awesome list.

    Jupyter Notebook
    Ver en GitHub↗19,404
  • meituan-automl/mobilevlmAvatar de Meituan-AutoML

    Meituan-AutoML/MobileVLM

    1,359Ver en GitHub↗

    MobileVLM: Vision Language Model for Mobile Devices

    Listed in the “Vision Language Model” section of the Ailia Models awesome list.

    Python
    Ver en GitHub↗1,359
  • tosiyuki/llava-jpT

    tosiyuki/LLaVA-JP

    0Ver en GitHub↗

    Listed in the “Vision Language Model” section of the Ailia Models awesome list.

    Ver en GitHub↗0
  1. Home
  2. Part of an Awesome List
  3. More to explore
  4. Vision Language Model