6 रिपॉजिटरी
Combining visual and textual data into a shared embedding space for multimodal reasoning.
Distinguishing note: The candidates focus on visualizers or QA benchmarks, not the core architectural integration of vision and language modalities.
Explore 6 awesome GitHub repositories matching artificial intelligence & ml · Visual-Language Multimodal Integration. Refine with filters or upvote what's useful.
This project provides a foundational framework and reference implementation for executing causal language modeling and multimodal reasoning on local systems. It includes a set of core components for managing model assets, a fine-tuning framework, and structural definitions required to instantiate transformer-based architectures. The system is distinguished by its ability to process combined text and image inputs through multimodal transformer models for visual reasoning and document analysis. It also supports the deployment of quantized models, reducing memory footprints through low-precision
Integrates visual and textual data streams into a shared embedding space to enable cross-modal reasoning.
CogVLM is a multimodal large language model designed to integrate visual and textual data for reasoning about images and generating natural language. It functions as a visual question answering system that analyzes image content to provide detailed descriptions or answer specific questions. The project includes a visual grounding model capable of mapping text descriptions to precise bounding box coordinates within an image. It also features a vision-based automation agent that analyzes screen captures to generate execution plans and interaction coordinates for software interfaces. The system
Integrates visual features and text embeddings into a shared space for reasoning across image and language inputs.
nanoVLM छोटे विज़न-लैंग्वेज मॉडल्स के लिए एक ट्रेनिंग फ्रेमवर्क और टूलकिट है। यह इमेज इनपुट को टेक्स्टुअल विवरणों के साथ जोड़ने और नेचुरल लैंग्वेज में उत्तर उत्पन्न करने के लिए PyTorch-आधारित वातावरण प्रदान करता है। इस प्रोजेक्ट में मॉडल वेट्स को सेव और लोड करने के लिए एक क्लाउड मॉडल वर्ज़निंग टूल शामिल है, जो विभिन्न एनवायरनमेंट में एसेट्स को सिंक्रोनाइज़ करता है। इसमें विज़न-लैंग्वेज मॉडल्स की सटीकता और विश्वसनीयता को मापने के लिए एक समर्पित इवैल्यूएशन सूट भी है। फ्रेमवर्क VRAM खपत माप के माध्यम से GPU रिसोर्स प्लानिंग को कवर करता है और चेकपॉइंट-आधारित स्टेट पर्सिस्टेंस के साथ ट्रेनिंग स्टेबिलिटी को मैनेज करता है।
Integrates visual feature extractors with language models in a shared embedding space for multimodal processing.
This project is a multi-label classification pipeline designed for genre prediction. It implements a machine learning workflow that assigns multiple category labels to a single item by processing both textual and visual input data. The system utilizes multimodal feature extraction to transform images and text descriptions into semantic vectors. This process includes using pre-trained networks for visual feature extraction and semantic word averaging for text analysis, allowing the model to integrate different data types into a unified input. The pipeline covers the full machine learning life
Combines visual features from images and semantic vectors from text into a unified input for genre prediction.
Spark NLP, Apache Spark वितरित कंप्यूटिंग फ्रेमवर्क पर निर्मित स्केलेबल टेक्स्ट विश्लेषण और मशीन लर्निंग के लिए एक टूलकिट है। यह बड़े पैमाने पर भाषाई डेटा को प्रोसेस करने के लिए एनोटेटर को अनुक्रमित करने के लिए एक मल्टीमॉडल मशीन लर्निंग फ्रेमवर्क और एक वितरित पाइपलाइन सिस्टम प्रदान करता है। लाइब्रेरी में प्रासंगिक वेक्टर एम्बेडिंग उत्पन्न करने के लिए एक ट्रांसफॉर्मर टेक्स्ट प्रोसेसर और बड़े भाषा मॉडल के प्रबंधन के लिए एक समर्पित अनुमान इंजन शामिल है। यह प्रोजेक्ट एक एकीकृत विज़न-भाषा आर्किटेक्चर के भीतर टेक्स्ट, ऑडियो और छवियों सहित विषम डेटा प्रकारों को प्रोसेस करने की अपनी क्षमता के माध्यम से खुद को अलग करता है। यह उन्नत जेनरेटिव AI क्षमताओं का समर्थन करता है जैसे कि प्रॉम्प्ट इंजीनियरिंग, प्रतिबंधित JSON आउटपुट के साथ संरचित एंटिटी निष्कर्षण, और नेटवर्क विलंबता को समाप्त करने के लिए स्थानीय अनुमान। इसके अतिरिक्त, यह टेक्स्ट और इमेज दोनों तौर-तरीकों में क्रॉस-भाषा अनुवाद और ज़ीरो-शॉट वर्गीकरण के लिए टूल प्रदान करता है। फ्रेमवर्क एंटिटी पहचान और भावना विश्लेषण के लिए पर्यवेक्षित मॉडल प्रशिक्षण, साथ ही निष्कर्षण प्रश्न उत्तर और दस्तावेज़ सारांश सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह समानता खोज के लिए वेक्टर डेटाबेस समर्थन को एकीकृत करता है और GPU त्वरण और केंद्रीकृत रजिस्ट्री के माध्यम से मॉडल लाइफसाइकिल प्रबंधन के लिए बुनियादी ढांचा प्रदान करता है। टूलकिट एक सार्वजनिक रिपॉजिटरी के माध्यम से कस्टम मॉडल और पाइपलाइनों के वितरण की अनुमति देता है और REST API के माध्यम से मॉडल की तैनाती का समर्थन करता है।
Combines visual and textual data into a shared embedding space for image captioning and document reasoning.
ml-mgie is a multimodal machine learning framework and image editor designed for instruction-based image manipulation. It utilizes multimodal large language models to translate natural language prompts into precise visual modifications, functioning as a text-to-image editing model. The system is a research implementation focused on aligning visual imagination with textual commands. It employs a training process based on image-pair datasets and descriptive instructions to learn how to execute complex visual edits. The framework covers capabilities in AI-powered visual content creation, includ
Integrates visual and textual encoders to interpret editing instructions and generate modification parameters.