15 repository-uri
Models capable of processing and generating across text, image, and other modalities.
Explore 15 awesome GitHub repositories matching part of an awesome list · Multimodal Models. Refine with filters or upvote what's useful.
MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im
Efficient multimodal model for visual and textual tasks.
Open-AutoGLM is an autonomous agent framework designed to perform complex user workflows on mobile devices. By translating natural language instructions into precise sequences of taps, scrolls, and text inputs, the system enables the automation of mobile application interactions and testing. The platform distinguishes itself through a combination of vision-language processing and reinforcement learning. It converts graphical user interfaces into structured data, allowing agents to parse screen elements and map natural language commands to coordinate-based actions. To ensure reliability, the s
Agentic multimodal model for automated device interaction.
DeepSeek-OCR is a vision processing framework designed to convert image-based text into machine-readable tokens for large language models. It functions as a document inference pipeline that encodes visual data into compact representations, enabling automated optical character recognition and document analysis workflows. The system distinguishes itself through a high-throughput architecture that utilizes hardware-accelerated batch inference to process large volumes of visual data. It incorporates dynamic resolution scaling to manage the balance between visual detail and token consumption, ensu
Specialized multimodal model for optical character recognition.
This project is a comprehensive framework and toolkit for developing, optimizing, and deploying transformer-based models across multimodal, document intelligence, and natural language processing tasks. It provides a unified neural architecture that processes text, vision, audio, and document layout data through a shared set of weights, enabling researchers and developers to build foundational models that align cross-modal representations. The platform distinguishes itself through advanced training and inference strategies designed for large-scale deep learning. It incorporates specialized mec
Framework for transformer-based optical character recognition.
Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw
Multimodal model supporting video and image-text processing.
:sparkles::sparkles:Latest Advances on Multimodal Large Language Models
Collection of papers and datasets for multimodal language models.
InternVL is a vision-language model framework that fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning. It provides a system for multimodal inference and dialogue, enabling the processing of images and text to answer questions or generate descriptions. The project is distinguished by its high-resolution image processing, which uses dynamic tiling to maintain detail for images up to 4K resolution, and its chain-of-thought visual reasoning for solving complex mathematical and spatial problems. It also supports temporal frame sampling
Large-scale multimodal model for visual and textual reasoning.
ImageBind is a multi-modal embedding model and joint representation learner that maps images, text, audio, and other modalities into a single shared vector space. It functions as a cross-modal retrieval framework designed to bind multiple sensory inputs into one cohesive mathematical embedding. The system uses a contrastive learning architecture to align disparate data types by maximizing the similarity between related samples. This allows the model to perform zero-shot multimodal classification and execute cross-modal data retrieval, such as locating visual content via natural language descr
Embedding space model for binding multiple data modalities.
Dolphin is a multimodal layout analyzer and image-to-structure converter that transforms photographed or digital document images into machine-readable structured data. It functions as an LLM document parser, utilizing vision-language models to simultaneously predict spatial layout and text content. The system is designed as a concurrent document processor, employing parallel document parsing to process multiple elements across distributed compute nodes. This high-throughput approach reduces the total time required to convert large volumes of images into structured formats. The project covers
Multimodal model for image and text integration.
ERNIE is a development toolkit for training, fine-tuning, and deploying large language models built on the PaddlePaddle deep learning platform. It provides a comprehensive suite of core components, including an inference server for vision and language models, a training and fine-tuning toolkit, and a framework for building retrieval-augmented generation systems using private knowledge bases. The project features multimodal AI models capable of reasoning across text, images, and video to perform complex visual understanding and information extraction. It distinguishes itself through specialize
Implements a model architecture capable of reasoning across text, images, and video for visual information extraction.
Qwen-Image is a text-to-image model and large language model image generation framework. It functions as an AI image editing suite and a personalized image trainer, capable of producing high-fidelity visuals and accurate typography from natural language descriptions. The system is distinguished by its precision text rendering engine, which integrates multi-script calligraphy and layout-coherent alphabetic text into images. It provides specialized capabilities for subject identity preservation and consistent subject generation across different poses and viewpoints, alongside a training pipelin
Multimodal model with advanced image understanding capabilities.
Pythia este un framework de cercetare multimodală și un sistem de antrenament distribuit conceput pentru construirea, antrenarea și evaluarea modelelor mari care combină date vizuale și lingvistice. Oferă un mediu modular pentru dezvoltarea modelelor vision-language, concentrându-se pe integrarea input-urilor de imagine și text în reprezentări de caracteristici partajate. Framework-ul utilizează o arhitectură modulară care decuplează blocurile de construcție ale modelului în componente interschimbabile, permițând configurarea flexibilă a modulelor de viziune și limbaj. Include o suită de benchmark-uri pentru executarea modelelor de referință pe seturi de date standardizate, pentru a stabili linii de bază de performanță consistente pentru sarcinile vision-language. Sistemul suportă pipeline-uri de antrenament distribuit pentru a scala dezvoltarea modelelor pe mai multe noduri de calcul și utilizează fișiere de configurare externe pentru maparea hiperparametrilor, pentru a asigura reproductibilitatea cercetării.
Provides a modular environment for building and training models capable of processing text and images.
Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
Multimodal model specialized for document and visual analysis.
This repo contains code to run models from our paper Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.
Unified model for diverse vision and language tasks.