awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

22 مستودعات

Awesome GitHub RepositoriesVisual Question Answering

Models and frameworks that answer natural language questions about visual content.

Distinct from Visual Question Answering Evaluation: Focuses on the actual task of answering questions, whereas the sibling focuses on the evaluation of those answers

Explore 22 awesome GitHub repositories matching artificial intelligence & ml · Visual Question Answering. Refine with filters or upvote what's useful.

Awesome Visual Question Answering GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • yujiangshui/a-programmers-guide-to-englishالصورة الرمزية لـ yujiangshui

    yujiangshui/A-Programmers-Guide-to-English

    16,428عرض على GitHub↗

    This project is a systematic framework for English language acquisition that applies structured workflows and cognitive strategies to build linguistic proficiency. It focuses on the construction of a linguistic knowledge base, enabling learners to master vocabulary and grammar through methodical training. The methodology is distinguished by its use of computer science concepts, such as mental-model-based learning and memory buffers, to organize progression. It emphasizes a cognitive-translation bypass to develop target language thinking, reducing mental latency by processing information direc

    Uses short question-answering exercises to build rapid auditory comprehension and reaction speed.

    englishenglish-learning
    عرض على GitHub↗16,428
  • salesforce/lavisالصورة الرمزية لـ salesforce

    salesforce/LAVIS

    11,236عرض على GitHub↗

    LAVIS is a multimodal large language model framework and vision-language model library. It provides tools for training and evaluating models that integrate visual, textual, and audio data, serving as a cross-modal feature extractor and a zero-shot visual reasoning engine. The framework distinguishes itself by using frozen-backbone integration, where pretrained encoders remain non-trainable while lightweight adapter layers are updated. It employs cross-modal feature alignment to map different representations into a shared embedding space and utilizes a modular model wrapper to swap vision and

    Responds to free-form natural language questions about image content using integrated multimodal models.

    Jupyter Notebook
    عرض على GitHub↗11,236
  • vikhyat/moondreamالصورة الرمزية لـ vikhyat

    vikhyat/moondream

    9,769عرض على GitHub↗

    Moondream is a small-scale vision language model designed to reason across images to generate captions and answer natural language questions. It functions as an edge-optimized system capable of performing visual question answering, image captioning, and object detection. The project distinguishes itself through a lightweight architecture designed for local inference on embedded devices, workstations, and air-gapped hardware. It supports the execution of models on local GPUs and Apple Silicon to ensure data privacy and low latency. The system's capabilities include identifying precise object

    Allows users to ask natural language questions about the contents of an image to extract context.

    Python
    عرض على GitHub↗9,769
  • intel/ipex-llmالصورة الرمزية لـ intel

    intel/ipex-llm

    8,836عرض على GitHub↗

    Intel XPU LLM Acceleration Library is a toolkit designed to accelerate large language model inference and finetuning on Intel CPUs, GPUs, and NPUs. It provides a distributed inference engine for scaling models across multiple accelerators, a multimodal model runtime for vision and speech tasks, and a low-bit model quantization tool for converting weights into INT4, FP8, and GGUF formats. The project features a parameter-efficient finetuning framework that enables model adaptation using QLoRA and DPO on Intel hardware. It distinguishes itself by providing specialized optimizations for Intel XP

    Predicts tokens based on combined image and text prompts to perform visual question answering.

    Python
    عرض على GitHub↗8,836
  • apple/ml-fastvlmالصورة الرمزية لـ apple

    apple/ml-fastvlm

    7,375عرض على GitHub↗

    This project is a vision language model framework and vision-to-text pipeline designed for deploying and optimizing models that process both images and text. It provides an on-device inference engine and a vision language model framework to run quantized models locally on mobile and desktop hardware accelerators. The framework features a model quantization toolkit to reduce weight precision for lower memory footprints and increased execution speed on specialized silicon. It also includes an efficient vision encoder utilizing a hybrid encoding system to compress image tokens, which reduces pro

    Implements on-device models and frameworks that answer natural language questions about visual content while maintaining privacy.

    Python
    عرض على GitHub↗7,375
  • thudm/cogvlmالصورة الرمزية لـ THUDM

    THUDM/CogVLM

    6,742عرض على GitHub↗

    CogVLM is a multimodal large language model designed to integrate visual and textual data for reasoning about images and generating natural language. It functions as a visual question answering system that analyzes image content to provide detailed descriptions or answer specific questions. The project includes a visual grounding model capable of mapping text descriptions to precise bounding box coordinates within an image. It also features a vision-based automation agent that analyzes screen captures to generate execution plans and interaction coordinates for software interfaces. The system

    Analyzes images and text to answer natural language questions and provide detailed visual descriptions.

    Python
    عرض على GitHub↗6,742
  • clovaai/donutالصورة الرمزية لـ clovaai

    clovaai/donut

    6,789عرض على GitHub↗

    Donut is an OCR-free document transformer and end-to-end document parser. It functions as a neural network that converts unstructured document images directly into structured data or text without the use of an external optical character recognition engine. The project includes a synthetic document generator to create artificial images and ground-truth labels for training. It employs a transformer model to perform visual question answering and document image classification based on visual layout and text. The system covers several document understanding capabilities, including structured info

    Produces text answers to natural language questions by analyzing the visual and spatial content of document images.

    Pythoncomputer-visiondocument-aieccv-2022
    عرض على GitHub↗6,789
  • firebase/genkitالصورة الرمزية لـ firebase

    firebase/genkit

    6,121عرض على GitHub↗

    Genkit is an open-source framework for building AI-powered applications. It provides a unified interface for connecting to hundreds of generative AI models from multiple providers, enabling text, image, audio, and video generation through a single API. The framework structures multi-step AI interactions—including chat, retrieval-augmented generation, tool use, and agentic workflows—as composable, traceable flows with built-in streaming and state management. The framework distinguishes itself through a comprehensive developer toolkit that includes a command-line interface and a local developer

    Transcribes speech, answers questions, or summarizes recordings from audio files.

    TypeScript
    عرض على GitHub↗6,121
  • sylinko/everywhereالصورة الرمزية لـ Sylinko

    Sylinko/Everywhere

    6,093عرض على GitHub↗

    Everywhere is a desktop AI assistant that understands whatever is on your screen and can act across applications without requiring screenshots or manual context switching. It reads structured UI data through accessibility and automation APIs to perceive the active application and visible content, then provides context-aware help, summaries, translations, and answers to natural language questions about what you are viewing. The tool distinguishes itself by combining on-screen content analysis with a multi-LLM agent platform that routes requests to providers like OpenAI, Anthropic, and local mo

    Responds to natural language queries by interpreting the captured screen context and providing relevant answers or actions.

    C#aiai-agentsai-assistant
    عرض على GitHub↗6,093
  • getstream/vision-agentsالصورة الرمزية لـ GetStream

    GetStream/Vision-Agents

    6,029عرض على GitHub↗

    Responds to natural-language questions about the content of video frames using a vision-language model.

    Pythonagentic-aiagentsai
    عرض على GitHub↗6,029
  • facebookresearch/mmfالصورة الرمزية لـ facebookresearch

    facebookresearch/mmf

    5,635عرض على GitHub↗

    MMF is a modular framework for building, training, and evaluating vision-and-language models. It provides a configuration-driven experiment system where model, dataset, and training parameters are defined through composable YAML files, alongside a curated model zoo of pretrained checkpoints for state-of-the-art multimodal architectures. The framework includes a multimodal dataset loader that downloads, processes, and batches vision-and-language data, and a vision-language model trainer supporting distributed training, mixed precision, and checkpoint-based resumption. The framework distinguish

    Processes visual questions by reading text in images and combining it with visual objects to predict answers.

    Pythoncaptioningdeep-learningdialog
    عرض على GitHub↗5,635
  • salesforce/blipالصورة الرمزية لـ salesforce

    salesforce/BLIP

    5,676عرض على GitHub↗

    BLIP is a vision-language model framework that combines contrastive, matching, and language modeling objectives to align images with text. Built on a multimodal encoder-decoder architecture, it supports distributed data-parallel training with cosine learning rate scheduling and sliding-window metric tracking for training stability. The framework provides capabilities for image captioning, visual question answering, and cross-modal retrieval, scoring semantic alignment between images and text through learned embeddings. It includes toolkits for fine-tuning pre-trained models on custom datasets

    Trains vision-language models to answer natural language questions about visual content.

    Jupyter Notebookimage-captioningimage-text-retrievalvision-and-language-pre-training
    عرض على GitHub↗5,676
  • deepseek-ai/deepseek-vl2الصورة الرمزية لـ deepseek-ai

    deepseek-ai/DeepSeek-VL2

    5,302عرض على GitHub↗

    DeepSeek-VL2 هو نموذج لغوي كبير متعدد الوسائط ونظام رؤية لغوية مصمم لتحليل المشاهد المرئية وتوليد نص وصفي. يعمل كنموذج للإجابة على الأسئلة المرئية والتأريض المرئي، وقادر على استخراج المعلومات من المستندات وتحديد كائنات أو مناطق محددة داخل الصور بناءً على أوصاف نصية. يستخدم المشروع معمارية خليط من الخبراء (mixture-of-experts) لمعالجة مدخلات الصور والنصوص المدمجة. تم تحسينه للاستدلال من خلال التعبئة التزايدية (incremental prefilling)، مما يقلل من متطلبات ذاكرة GPU على الأجهزة. يغطي النموذج تحليل البيانات متعدد الوسائط وفهم المستندات المرئية، بما في ذلك تفسير المخططات والتخطيطات. يقوم بإجراء استدلال مرئي وتأريض لمطابقة الاستعلامات النصية مع المحتوى المرئي المقابل.

    Extracts information from images and documents to answer complex natural language queries.

    Python
    عرض على GitHub↗5,302
  • huggingface/nanovlmالصورة الرمزية لـ huggingface

    huggingface/nanoVLM

    4,917عرض على GitHub↗

    nanoVLM is a training framework and toolkit for small vision-language models. It provides a PyTorch-based environment for training and fine-tuning models to associate image inputs with textual descriptions and generate natural language answers. The project includes a cloud model versioning tool for saving and loading model weights to centralized repositories to synchronize assets across environments. It also features a dedicated evaluation suite for measuring the accuracy and reliability of vision-language models against standard task datasets. The framework covers GPU resource planning thro

    Generates natural language answers and descriptive captions based on visual content.

    Python
    عرض على GitHub↗4,917
  • tencentcloudadp/youtu-agentالصورة الرمزية لـ TencentCloudADP

    TencentCloudADP/youtu-agent

    4,576عرض على GitHub↗

    Youtu Agent is an open-source framework for building, running, and evaluating autonomous agents powered by large language models. It provides the core infrastructure for creating agents that follow reasoning loops, use toolkits, and coordinate with other agents to solve complex tasks, all managed through YAML-driven configuration files. The framework distinguishes itself through its support for multi-agent orchestration, where a planner agent decomposes tasks and coordinates specialized worker agents, and through its integration with the Model Context Protocol for connecting to external toolk

    Answer a user's natural-language question about the content of a given image by querying a vision-language model.

    Pythonagent-frameworkagentsopenai-agents
    عرض على GitHub↗4,576
  • moonshotai/kimi-audioالصورة الرمزية لـ MoonshotAI

    MoonshotAI/Kimi-Audio

    4,492عرض على GitHub↗

    Kimi-Audio is a large language model audio foundation model designed to understand audio input and generate high-fidelity speech responses in real time. It functions as a unified system encompassing a text-to-speech synthesis engine and a speech-to-text transcription tool. The project enables real-time audio conversations through a multi-modal conversation loop and chunk-wise streaming detokenization to reduce playback latency. It provides controls over speech speed, accent, and emotional tone during conversational audio generation. The system covers audio intelligence capabilities, includin

    Responds to natural-language queries about the content of an audio clip, such as identifying sounds or answering factual questions.

    Python
    عرض على GitHub↗4,492
  • deepseek-ai/deepseek-vlالصورة الرمزية لـ deepseek-ai

    deepseek-ai/DeepSeek-VL

    4,134عرض على GitHub↗

    DeepSeek-VL هو نموذج لغة كبير متعدد الوسائط ومحرك استدلال من الصورة إلى النص. يعمل كنموذج رؤية-لغة ونظام للإجابة على الأسئلة المرئية يدمج الإدراك البصري مع الاستدلال اللغوي لفهم ووصف الصور. يمكن المشروع من فهم الصور متعددة الوسائط وتحليل صور المستندات، وتحديداً معالجة لقطات شاشة صفحات الويب والمخططات التقنية. ويوفر إمكانيات للذكاء الاصطناعي التحادثي المرئي، مما يسمح للمستخدمين بالتفاعل مع البيانات المرئية لاستخراج الرؤى وإجراء استدلال معقد عبر أنواع مختلفة من المعلومات المرئية. يستخدم النظام بنية محول رؤية-لغة تجمع بين محول رؤية (Vision Transformer) للترميز البصري ونموذج لغة كبير. ويستخدم ضبط التعليمات متعدد الوسائط وطبقة إسقاط لمحاذاة متجهات الميزات المرئية مع مساحة تضمين نموذج اللغة لتوليد النص التلقائي.

    Provides a system for answering natural language questions about the contents and details of visual data.

    Python
    عرض على GitHub↗4,134
  • johnsnowlabs/spark-nlpالصورة الرمزية لـ JohnSnowLabs

    JohnSnowLabs/spark-nlp

    4,135عرض على GitHub↗

    Spark NLP هي مجموعة أدوات لتحليل النصوص القابل للتوسع والتعلم الآلي مبنية على إطار عمل الحوسبة الموزعة Apache Spark. توفر إطار عمل للتعلم الآلي متعدد الوسائط ونظام خط أنابيب موزع لتسلسل أدوات التعليق لمعالجة البيانات اللغوية على نطاق واسع. تتضمن المكتبة معالج نصوص محولاً (transformer) لتوليد تضمينات متجهات سياقية ومحرك استدلال مخصص لإدارة نماذج اللغة الكبيرة. يتميز المشروع بقدرته على معالجة أنواع البيانات غير المتجانسة، بما في ذلك النصوص والصوت والصور، ضمن بنية رؤية-لغة موحدة. ويدعم إمكانيات الذكاء الاصطناعي التوليدي المتقدمة مثل هندسة الأوامر (prompt engineering)، واستخراج الكيانات المهيكلة مع مخرجات JSON مقيدة، والاستدلال المحلي للقضاء على زمن انتقال الشبكة. بالإضافة إلى ذلك، يوفر أدوات للترجمة عبر اللغات والتصنيف بدون تدريب عبر كل من وسائط النص والصورة. يغطي إطار العمل مجموعة واسعة من الإمكانيات، بما في ذلك تدريب النماذج الخاضعة للإشراف للتعرف على الكيانات وتحليل المشاعر، بالإضافة إلى الإجابة على الأسئلة الاستخراجية وتلخيص المستندات. ويدمج دعم قاعدة بيانات المتجهات للبحث عن التشابه ويوفر بنية تحتية لتسريع GPU وإدارة دورة حياة النموذج عبر سجل مركزي. تسمح مجموعة الأدوات بتوزيع النماذج وخطوط الأنابيب المخصصة عبر مستودع عام وتدعم نشر النماذج عبر واجهات برمجة تطبيقات REST.

    Generates text answers to natural language questions about an input image by merging vision and text embeddings.

    Scala
    عرض على GitHub↗4,135
  • huggingface/smollmالصورة الرمزية لـ huggingface

    huggingface/smollm

    3,624عرض على GitHub↗

    SmolLM is a project dedicated to the development of small language models. It focuses on training and fine-tuning compact models that maintain high performance while utilizing fewer parameters. The project emphasizes efficient AI inference and on-device text generation, aiming to enable the deployment of lightweight models on edge devices with limited memory and processing power. It utilizes synthetic data generation to produce artificial datasets that improve the reasoning and training of these AI systems. The system supports a variety of optimization and training capabilities, including we

    Interprets multiple images and text in a single conversation to perform visual question answering.

    Python
    عرض على GitHub↗3,624
  • nunchaku-ai/comfyui-nunchakuالصورة الرمزية لـ nunchaku-ai

    nunchaku-ai/ComfyUI-nunchaku

    2,901عرض على GitHub↗

    ComfyUI-nunchaku is a 4-bit diffusion inference engine and a set of nodes for running low-precision quantized diffusion models within ComfyUI visual workflows. It provides a backend that reduces memory overhead and increases generation speed for transformer models. The project includes specialized tools for identity-preserving generation and an image-to-image guidance toolkit that uses depth maps and reference images. It also features a multimodal visual question answering implementation and a utility for merging multiple quantized model files into single unified files. The engine covers a b

    Provides a visual question answering implementation that processes images and text using quantized multimodal models.

    Pythoncomfyuidiffusionflux
    عرض على GitHub↗2,901
السابق12التالي
  1. Home
  2. Artificial Intelligence & ML
  3. Visual Question Answering

استكشف الوسوم الفرعية

  • Audio Question Answering1 وسم فرعيResponds to natural-language queries about audio content, such as identifying sounds or answering factual questions. **Distinct from Visual Question Answering:** Distinct from Visual Question Answering: focuses on answering questions about audio content, not visual content.
  • On-Screen Content Question AnswerersAnswers natural language questions by interpreting captured screen context from any application. **Distinct from Visual Question Answering:** Distinct from Visual Question Answering: answers questions about structured UI data and text on screen, not images or visual scenes.