4 repositorios
Explore 4 awesome GitHub repositories matching part of an awesome list · Vision Language Model. Refine with filters or upvote what's useful.
LLaVA is a multimodal large language model architecture designed to process and interpret both image and text inputs to generate natural language responses. It functions as a research-oriented platform for visual instruction tuning, providing a framework to align language models with human intent through training on diverse datasets of paired images and text queries. The system distinguishes itself through a specialized vision-language training pipeline that connects visual data to language models using projection layers and instruction-based fine-tuning. It supports distributed inference by
Listed in the “Vision Language Model” section of the Ailia Models awesome list.
Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw
Listed in the “Vision Language Model” section of the Ailia Models awesome list.
MobileVLM: Vision Language Model for Mobile Devices
Listed in the “Vision Language Model” section of the Ailia Models awesome list.
Listed in the “Vision Language Model” section of the Ailia Models awesome list.