5 रिपॉजिटरी
Systems that integrate visual and language data to perform complex reasoning and maintain context across visual inputs.
Distinct from Multimodal Understanding: The existing candidates focus on document-specific or audio-specific understanding rather than general multimodal visual reasoning.
Explore 5 awesome GitHub repositories matching artificial intelligence & ml · Multimodal Visual Understanding. Refine with filters or upvote what's useful.
CogVLM is a multimodal large language model designed for visual reasoning and multi-turn dialogue. It functions as a visual grounding model and a quantized vision model, combining text and image processing to perform complex understanding and maintain context across visual inputs. The project includes capabilities as a GUI automation agent, allowing it to analyze application screenshots, plan operational steps, and return precise screen coordinates for interface interaction. It further supports visual grounding by generating bounding box coordinates to map text descriptions to specific spatia
Integrates visual and language data to perform complex understanding and maintain context across visual inputs.
DeepSeek-VL2 is a multimodal large language model and vision-language system designed to analyze visual scenes and generate descriptive text. It functions as a visual question answering and visual grounding model, capable of extracting information from documents and locating specific objects or regions within images based on textual descriptions. The project utilizes a mixture-of-experts architecture to process combined image and text inputs. It is optimized for inference through incremental prefilling, which reduces the GPU memory requirements on hardware. The model covers multimodal data a
Processes combined image and text inputs to perform complex multimodal visual reasoning.
This project is a computer vision dataset and image annotation repository designed for training and evaluating machine learning models. It provides a large collection of labeled images, serving as an object detection benchmark and a source of pixel-level segmentation data. The repository distinguishes itself as a multimodal visual dataset by pairing images with synchronized voice, text, and mouse traces to support narrative understanding. It further enables the analysis of model fairness through the inclusion of demographic attributes and exhaustive annotations. The dataset covers a broad ra
Integrates visual and language data, linking voice traces and narratives to image regions for complex reasoning.
DeepSeek-VL एक मल्टीमॉडल बड़ा भाषा मॉडल और इमेज-टू-टेक्स्ट रीजनिंग इंजन है। यह एक विज़न-भाषा मॉडल और विज़ुअल प्रश्न उत्तर प्रणाली के रूप में कार्य करता है जो छवियों को समझने और उनका वर्णन करने के लिए भाषाई तर्क के साथ विज़ुअल धारणा को एकीकृत करता है। प्रोजेक्ट मल्टीमॉडल इमेज समझ और दस्तावेज़ इमेज विश्लेषण को सक्षम बनाता है, विशेष रूप से वेब पेजों और तकनीकी आरेखों के स्क्रीनशॉट को प्रोसेस करता है। यह विज़ुअल संवादात्मक AI के लिए क्षमताएं प्रदान करता है, जिससे उपयोगकर्ता अंतर्दृष्टि निकालने और विभिन्न प्रकार की विज़ुअल जानकारी में जटिल तर्क करने के लिए विज़ुअल डेटा के साथ बातचीत कर सकते हैं। सिस्टम एक विज़न-भाषा ट्रांसफॉर्मर आर्किटेक्चर का उपयोग करता है जो विज़ुअल एन्कोडिंग के लिए एक विज़न ट्रांसफॉर्मर को एक बड़े भाषा मॉडल के साथ जोड़ता है। यह ऑटोरिग्र्रेसिव टेक्स्ट जनरेशन के लिए भाषा मॉडल के एम्बेडिंग स्पेस के साथ विज़ुअल फ़ीचर वैक्टर को संरेखित करने के लिए मल्टीमॉडल इंस्ट्रक्शन ट्यूनिंग और एक प्रोजेक्शन लेयर का उपयोग करता है।
Integrates visual and language data to analyze images and diagrams through complex multimodal reasoning.
Kimi-code is a command-line interface and orchestration framework designed to integrate autonomous AI agents into software development workflows. It functions as a terminal-based assistant that manages multi-step coding tasks, including planning, file system modifications, shell command execution, and test running, all while maintaining conversational context within a local development environment. The project distinguishes itself through a focus on secure, autonomous agent orchestration and granular control over AI interactions. It enforces strict security by requiring explicit user approval
Analyzes screen recordings and video clips alongside text prompts to understand visual context.