awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

25 रिपॉजिटरी

Awesome GitHub RepositoriesVision-Language Models

Architectures and resources for models integrating visual and linguistic processing.

Explore 25 awesome GitHub repositories matching artificial intelligence & ml · Vision-Language Models. Refine with filters or upvote what's useful.

Awesome Vision-Language Models GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • hiyouga/llama-factoryhiyouga का अवतार

    hiyouga/LLaMA-Factory

    72,241GitHub पर देखें↗

    LLaMA-Factory is a comprehensive suite for dataset preparation, model fine-tuning, memory optimization, and standardized API deployment. It provides a unified platform for the supervised and reward-based fine-tuning of large language models and vision-language models. The framework includes a specialized toolkit for training vision-language models and a model serving interface that deploys trained models through high-performance APIs. It utilizes precision tuning and quantization techniques to reduce the hardware requirements and memory footprint of large models. The system covers data pipel

    Provides specialized capabilities for training vision-language models using supervised and reward-based methods.

    Python
    GitHub पर देखें↗72,241
  • datawhalechina/hello-agentsdatawhalechina का अवतार

    datawhalechina/hello-agents

    59,685GitHub पर देखें↗

    This project provides a comprehensive framework for building, training, and managing autonomous agents. It enables the construction of systems that utilize language models to plan, manage memory, and execute multi-step tasks through iterative reasoning loops and tool-based actions. The framework distinguishes itself by offering specialized capabilities for interacting with graphical user interfaces and legacy software, allowing agents to perceive visual elements and perform actions like a human user. It supports complex, cross-application workflows through graph-based orchestration and provid

    Links visual encoders with language models using cross-attention mechanisms to enable accurate multimodal understanding.

    Pythonagentllmrag
    GitHub पर देखें↗59,685
  • chenfei-wu/taskmatrixchenfei-wu का अवतार

    chenfei-wu/TaskMatrix

    34,082GitHub पर देखें↗

    TaskMatrix is a multimodal AI chat interface and visual task orchestrator. It combines language models with visual recognition to enable the exchange, analysis, and modification of images within a conversational environment. The system coordinates multiple foundation models through orchestration pipelines that chain language, detection, and segmentation models. This allows for complex visual operations, such as using text instructions to guide image masking and executing modular inpainting workflows to edit specific image regions. The project includes a computer vision toolset for object det

    Integrates vision-language models to reason about image content within a conversational chat interface.

    Python
    GitHub पर देखें↗34,082
  • sgl-project/sglangsgl-project का अवतार

    sgl-project/sglang

    29,079GitHub पर देखें↗

    Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr

    Processes visual inputs using raw images or precomputed embeddings to generate text responses from multimodal models.

    Pythonattentionblackwellcuda
    GitHub पर देखें↗29,079
  • vision-cair/minigpt-4Vision-CAIR का अवतार

    Vision-CAIR/MiniGPT-4

    25,679GitHub पर देखें↗

    MiniGPT-4 is a multimodal AI framework and large language model that integrates vision encoders with language models to process and reason about combined image and text inputs. It functions as a vision-language model capable of image-based conversational AI, visual question answering, and multimodal logical reasoning. The project utilizes a pretrained vision-language integration strategy that connects a vision encoder to a language model via a linear projection layer. This approach employs frozen-backbone training to align visual representations with linguistic tokens while keeping the primar

    Integrates a vision encoder with a large language model to reason and converse about images.

    Python
    GitHub पर देखें↗25,679
  • openbmb/minicpm-vOpenBMB का अवतार

    OpenBMB/MiniCPM-V

    25,653GitHub पर देखें↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Analyzes high-resolution images and high-frame-rate video to generate descriptive text outputs.

    Python
    GitHub पर देखें↗25,653
  • qwenlm/qwen2-vlQwenLM का अवतार

    QwenLM/Qwen2-VL

    19,404GitHub पर देखें↗

    Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw

    Integrates visual and linguistic processing to perform object detection, document parsing, and spatial reasoning.

    Jupyter Notebook
    GitHub पर देखें↗19,404
  • qwenlm/qwen3-vlQwenLM का अवतार

    QwenLM/Qwen3-VL

    18,329GitHub पर देखें↗

    Qwen3-VL is a multimodal vision-language model designed to process and reason across images, videos, and text. It functions as a computer vision framework capable of identifying objects, extracting structured data from documents, and interpreting spatial elements within visual media. The system operates as an automated user interface interaction agent, interpreting screen data to navigate software and mobile applications. By utilizing a unified transformer architecture, it performs complex visual reasoning to execute user-defined tasks without manual input. Beyond interface navigation, the m

    Processes and reasons across images, videos, and text using a multimodal artificial intelligence architecture.

    Jupyter Notebook
    GitHub पर देखें↗18,329
  • xming521/weclonexming521 का अवतार

    xming521/WeClone

    18,028GitHub पर देखें↗

    WeClone is an end-to-end framework designed for the creation, training, and deployment of personalized conversational AI digital twins. By fine-tuning large language models on individual chat history, the platform enables the replication of unique communication styles, speech patterns, and conversational habits. The system manages the entire lifecycle of these digital avatars, from initial data preparation to final integration into messaging platforms for real-time interaction. The platform distinguishes itself through a comprehensive suite of data processing utilities that prepare raw messag

    Converts image content into descriptive text tokens to allow language models to interpret visual data during training.

    Pythonchat-historydigital-avatarllm
    GitHub पर देखें↗18,028
  • allenai/olmocrallenai का अवतार

    allenai/olmocr

    17,396GitHub पर देखें↗

    Olmocr is a distributed document processing framework designed to convert PDF and image files into structured markdown. It functions as a vision-based document parser that utilizes multimodal neural networks to interpret complex visual layouts and translate them into standardized text representations. The system operates as a remote inference orchestrator, offloading heavy document analysis tasks to external servers or cloud APIs to minimize local computational requirements. By employing a stateless worker architecture, it decouples document ingestion from inference, allowing for the distribu

    Utilizes multimodal neural networks to interpret complex visual document layouts and translate them into text.

    Python
    GitHub पर देखें↗17,396
  • sawyerhood/draw-a-uiSawyerHood का अवतार

    SawyerHood/draw-a-ui

    13,602GitHub पर देखें↗

    draw-a-ui is an AI vision UI generator and sketch-to-code tool that transforms hand-drawn sketches and digital wireframes into functional HTML and CSS. It serves as a mockup-to-HTML converter that interprets user interface layouts from images to produce corresponding web markup. The system utilizes vision-capable language models to automate the transition from visual design to web code. It employs a multimodal inference loop to process canvas snapshots and natural language instructions, generating structural layouts and responsive grid systems without the need for pre-defined component templa

    Integrates vision-language models to process combined image and text inputs for structural web code generation.

    TypeScriptaigptopenai
    GitHub पर देखें↗13,602
  • physical-intelligence/openpiPhysical-Intelligence का अवतार

    Physical-Intelligence/openpi

    12,377GitHub पर देखें↗

    OpenPi is a vision-language-action robot control framework designed to generate physical control actions for robotic systems. It functions as a distributed robot model trainer, a model format converter, and a robot action streaming server. The framework provides tools for transforming model checkpoints between different framework formats to ensure interoperability across various development environments. It also includes a server that uses websocket connections to stream model-generated control actions from remote inference servers to physical robot hardware in real-time. The system supports

    Generates direct physical control commands for robots by processing visual and textual inputs through VLA models.

    Python
    GitHub पर देखें↗12,377
  • codexu/note-gencodexu का अवतार

    codexu/note-gen

    12,173GitHub पर देखें↗

    Note-gen is an artificial intelligence-assisted note-taking application and knowledge management tool designed for local-first data ownership. It functions as a workspace that leverages language models to organize, summarize, and synthesize personal notes into structured documents while maintaining offline accessibility. The platform distinguishes itself through a multimodal workflow orchestrator that chains sequences of tasks to process text, images, and external data. By integrating vision-language models, it extracts information from visual inputs like screenshots and documents, converting

    Integrates multimodal models to process and convert visual data into structured text.

    TypeScriptagentchatbotknowledge-base
    GitHub पर देखें↗12,173
  • kornia/korniakornia का अवतार

    kornia/kornia

    11,238GitHub पर देखें↗

    Kornia is a differentiable computer vision library and cross-framework tensor vision toolset. It implements vision operations as differentiable tensors to enable integration into deep learning pipelines and supports the transpilation of operations across PyTorch, TensorFlow, JAX, and NumPy. The project provides specialized toolsets for geometric vision and stereo depth, including algorithms for 3D scene reconstruction, camera calibration, and pose estimation. It further distinguishes itself as a differentiable image augmentation framework, applying random geometric and color transformations w

    Combines computer vision operations with large language models to process both images and text.

    Pythonartificial-intelligencecomputer-visiondeep-learning
    GitHub पर देखें↗11,238
  • salesforce/lavissalesforce का अवतार

    salesforce/LAVIS

    11,236GitHub पर देखें↗

    LAVIS is a multimodal large language model framework and vision-language model library. It provides tools for training and evaluating models that integrate visual, textual, and audio data, serving as a cross-modal feature extractor and a zero-shot visual reasoning engine. The framework distinguishes itself by using frozen-backbone integration, where pretrained encoders remain non-trainable while lightweight adapter layers are updated. It employs cross-modal feature alignment to map different representations into a shared embedding space and utilizes a modular model wrapper to swap vision and

    Provides tools to merge visual and textual data to develop models capable of processing both modalities simultaneously.

    Jupyter Notebook
    GitHub पर देखें↗11,236
  • openvinotoolkit/openvinoopenvinotoolkit का अवतार

    openvinotoolkit/openvino

    10,414GitHub पर देखें↗

    OpenVINO is an AI inference engine and model serving platform designed to execute optimized deep learning models across CPUs, GPUs, and NPUs through a unified API. It includes a model optimization toolkit for converting, quantizing, and compressing models from various frameworks, alongside a specialized generative AI runtime for large language models. The project distinguishes itself through a plugin-based hardware acceleration layer that maps neural network operations to vendor-specific drivers. It features advanced execution mechanisms such as continuous batching, speculative decoding, and

    Deploys multimodal models capable of analyzing combined text and image inputs for visual reasoning.

    C++aicomputer-visiondeep-learning
    GitHub पर देखें↗10,414
  • opengvlab/internvlOpenGVLab का अवतार

    OpenGVLab/InternVL

    10,061GitHub पर देखें↗

    InternVL is a vision-language model framework that fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning. It provides a system for multimodal inference and dialogue, enabling the processing of images and text to answer questions or generate descriptions. The project is distinguished by its high-resolution image processing, which uses dynamic tiling to maintain detail for images up to 4K resolution, and its chain-of-thought visual reasoning for solving complex mathematical and spatial problems. It also supports temporal frame sampling

    Fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning.

    Pythongptgpt-4ogpt-4v
    GitHub पर देखें↗10,061
  • ailab-cvc/yolo-worldAILab-CVC का अवतार

    AILab-CVC/YOLO-World

    6,425GitHub पर देखें↗

    YOLO-World is a vision-language framework and open-vocabulary object detection model. It identifies objects in images and video based on free-form text prompts without requiring predefined category labels. The system enables the identification of arbitrary objects by fusing image features with text embeddings. It includes a specialized tool for automated image labeling, which generates bounding box annotations for custom datasets using text-based prompts. The project provides a deployment pipeline for converting models into quantized ONNX and TFLite formats, supporting real-time inference on

    Utilizes a vision-language model architecture that fuses image features with text embeddings for object recognition.

    Python
    GitHub पर देखें↗6,425
  • nvidia/isaac-gr00tNVIDIA का अवतार

    NVIDIA/Isaac-GR00T

    6,222GitHub पर देखें↗

    Processes multimodal inputs including natural language and camera images to generate motor commands for generalized robot manipulation.

    Jupyter Notebook
    GitHub पर देखें↗6,222
  • bytedance-seed/bagelByteDance-Seed का अवतार

    ByteDance-Seed/Bagel

    5,681GitHub पर देखें↗

    Processes and generates images within a single framework, handling visual question answering and text-to-image creation.

    Python
    GitHub पर देखें↗5,681
पिछला12अगला
  1. Home
  2. Artificial Intelligence & ML
  3. Machine Learning
  4. Multimodal Processing Tools
  5. Vision-Language Models

सब-टैग एक्सप्लोर करें

  • Action InferenceThe process of generating physical control commands from multimodal vision-language model outputs. **Distinct from Vision-Language Models:** Specifically targets the generation of physical actions, whereas Vision-Language Models covers the general architecture.
  • Action Output ModelsProcesses multimodal inputs including natural language and camera images to generate motor commands for generalized robot manipulation. **Distinct from Vision-Language Models:** Distinct from Vision-Language Models: extends vision-language processing to generate motor commands (action output), not just visual and linguistic understanding.