2 مستودعات
Frameworks that enable models to reason over visual content to answer natural language questions.
Distinct from Visual Question Answering Evaluation: Focuses on the reasoning framework for answering questions, whereas the parent focuses on the evaluation metrics
Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Reasoning Frameworks. Refine with filters or upvote what's useful.
LAVIS is a multimodal large language model framework and vision-language model library. It provides tools for training and evaluating models that integrate visual, textual, and audio data, serving as a cross-modal feature extractor and a zero-shot visual reasoning engine. The framework distinguishes itself by using frozen-backbone integration, where pretrained encoders remain non-trainable while lightweight adapter layers are updated. It employs cross-modal feature alignment to map different representations into a shared embedding space and utilizes a modular model wrapper to swap vision and
Leverages pretrained vision and language models within zero-shot frameworks to answer questions about visual content.
Qwen3-Omni is an omni-modal large language model designed to process and generate text, audio, images, and video within a single unified neural architecture. It functions as a real-time voice assistant and multimodal AI agent capable of reasoning across different media types and executing external tool-calling functions via APIs. The system supports low-latency conversational AI through autoregressive token streaming and natural turn-taking. It enables multilingual speech translation and generation across dozens of languages, featuring customizable speaker profiles and tones. The model's cap
Answers complex questions by aligning timing between audio and video streams to understand scenarios.