awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

4 Repos

Awesome GitHub RepositoriesVision Language Model

Explore 4 awesome GitHub repositories matching part of an awesome list · Vision Language Model. Refine with filters or upvote what's useful.

Awesome Vision Language Model GitHub Repositories

Finde die besten Repos mit KI.Wir suchen mit KI nach den am besten passenden Repositories.
  • haotian-liu/llavaAvatar von haotian-liu

    haotian-liu/LLaVA

    24,465Auf GitHub ansehen↗

    LLaVA is a multimodal large language model architecture designed to process and interpret both image and text inputs to generate natural language responses. It functions as a research-oriented platform for visual instruction tuning, providing a framework to align language models with human intent through training on diverse datasets of paired images and text queries. The system distinguishes itself through a specialized vision-language training pipeline that connects visual data to language models using projection layers and instruction-based fine-tuning. It supports distributed inference by

    Listed in the “Vision Language Model” section of the Ailia Models awesome list.

    Pythonchatbotchatgptfoundation-models
    Auf GitHub ansehen↗24,465
  • qwenlm/qwen2-vlAvatar von QwenLM

    QwenLM/Qwen2-VL

    19,404Auf GitHub ansehen↗

    Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw

    Listed in the “Vision Language Model” section of the Ailia Models awesome list.

    Jupyter Notebook
    Auf GitHub ansehen↗19,404
  • meituan-automl/mobilevlmAvatar von Meituan-AutoML

    Meituan-AutoML/MobileVLM

    1,359Auf GitHub ansehen↗

    MobileVLM: Vision Language Model for Mobile Devices

    Listed in the “Vision Language Model” section of the Ailia Models awesome list.

    Python
    Auf GitHub ansehen↗1,359
  • tosiyuki/llava-jpT

    tosiyuki/LLaVA-JP

    0Auf GitHub ansehen↗

    Listed in the “Vision Language Model” section of the Ailia Models awesome list.

    Auf GitHub ansehen↗0
  1. Home
  2. Part of an Awesome List
  3. More to explore
  4. Vision Language Model