awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

6 रिपॉजिटरी

Awesome GitHub RepositoriesVisual-Language Multimodal Integration

Combining visual and textual data into a shared embedding space for multimodal reasoning.

Distinguishing note: The candidates focus on visualizers or QA benchmarks, not the core architectural integration of vision and language modalities.

Explore 6 awesome GitHub repositories matching artificial intelligence & ml · Visual-Language Multimodal Integration. Refine with filters or upvote what's useful.

Awesome Visual-Language Multimodal Integration GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • meta-llama/llama-modelsmeta-llama का अवतार

    meta-llama/llama-models

    7,643GitHub पर देखें↗

    This project provides a foundational framework and reference implementation for executing causal language modeling and multimodal reasoning on local systems. It includes a set of core components for managing model assets, a fine-tuning framework, and structural definitions required to instantiate transformer-based architectures. The system is distinguished by its ability to process combined text and image inputs through multimodal transformer models for visual reasoning and document analysis. It also supports the deployment of quantized models, reducing memory footprints through low-precision

    Integrates visual and textual data streams into a shared embedding space to enable cross-modal reasoning.

    Python
    GitHub पर देखें↗7,643
  • thudm/cogvlmTHUDM का अवतार

    THUDM/CogVLM

    6,742GitHub पर देखें↗

    CogVLM is a multimodal large language model designed to integrate visual and textual data for reasoning about images and generating natural language. It functions as a visual question answering system that analyzes image content to provide detailed descriptions or answer specific questions. The project includes a visual grounding model capable of mapping text descriptions to precise bounding box coordinates within an image. It also features a vision-based automation agent that analyzes screen captures to generate execution plans and interaction coordinates for software interfaces. The system

    Integrates visual features and text embeddings into a shared space for reasoning across image and language inputs.

    Python
    GitHub पर देखें↗6,742
  • huggingface/nanovlmhuggingface का अवतार

    huggingface/nanoVLM

    4,917GitHub पर देखें↗

    nanoVLM छोटे विज़न-लैंग्वेज मॉडल्स के लिए एक ट्रेनिंग फ्रेमवर्क और टूलकिट है। यह इमेज इनपुट को टेक्स्टुअल विवरणों के साथ जोड़ने और नेचुरल लैंग्वेज में उत्तर उत्पन्न करने के लिए PyTorch-आधारित वातावरण प्रदान करता है। इस प्रोजेक्ट में मॉडल वेट्स को सेव और लोड करने के लिए एक क्लाउड मॉडल वर्ज़निंग टूल शामिल है, जो विभिन्न एनवायरनमेंट में एसेट्स को सिंक्रोनाइज़ करता है। इसमें विज़न-लैंग्वेज मॉडल्स की सटीकता और विश्वसनीयता को मापने के लिए एक समर्पित इवैल्यूएशन सूट भी है। फ्रेमवर्क VRAM खपत माप के माध्यम से GPU रिसोर्स प्लानिंग को कवर करता है और चेकपॉइंट-आधारित स्टेट पर्सिस्टेंस के साथ ट्रेनिंग स्टेबिलिटी को मैनेज करता है।

    Integrates visual feature extractors with language models in a shared embedding space for multimodal processing.

    Python
    GitHub पर देखें↗4,917
  • spandan-madan/deeplearningprojectSpandan-Madan का अवतार

    Spandan-Madan/DeepLearningProject

    4,785GitHub पर देखें↗

    This project is a multi-label classification pipeline designed for genre prediction. It implements a machine learning workflow that assigns multiple category labels to a single item by processing both textual and visual input data. The system utilizes multimodal feature extraction to transform images and text descriptions into semantic vectors. This process includes using pre-trained networks for visual feature extraction and semantic word averaging for text analysis, allowing the model to integrate different data types into a unified input. The pipeline covers the full machine learning life

    Combines visual features from images and semantic vectors from text into a unified input for genre prediction.

    HTMLdeep-learningmachine-learningneural-networks
    GitHub पर देखें↗4,785
  • johnsnowlabs/spark-nlpJohnSnowLabs का अवतार

    JohnSnowLabs/spark-nlp

    4,135GitHub पर देखें↗

    Spark NLP, Apache Spark वितरित कंप्यूटिंग फ्रेमवर्क पर निर्मित स्केलेबल टेक्स्ट विश्लेषण और मशीन लर्निंग के लिए एक टूलकिट है। यह बड़े पैमाने पर भाषाई डेटा को प्रोसेस करने के लिए एनोटेटर को अनुक्रमित करने के लिए एक मल्टीमॉडल मशीन लर्निंग फ्रेमवर्क और एक वितरित पाइपलाइन सिस्टम प्रदान करता है। लाइब्रेरी में प्रासंगिक वेक्टर एम्बेडिंग उत्पन्न करने के लिए एक ट्रांसफॉर्मर टेक्स्ट प्रोसेसर और बड़े भाषा मॉडल के प्रबंधन के लिए एक समर्पित अनुमान इंजन शामिल है। यह प्रोजेक्ट एक एकीकृत विज़न-भाषा आर्किटेक्चर के भीतर टेक्स्ट, ऑडियो और छवियों सहित विषम डेटा प्रकारों को प्रोसेस करने की अपनी क्षमता के माध्यम से खुद को अलग करता है। यह उन्नत जेनरेटिव AI क्षमताओं का समर्थन करता है जैसे कि प्रॉम्प्ट इंजीनियरिंग, प्रतिबंधित JSON आउटपुट के साथ संरचित एंटिटी निष्कर्षण, और नेटवर्क विलंबता को समाप्त करने के लिए स्थानीय अनुमान। इसके अतिरिक्त, यह टेक्स्ट और इमेज दोनों तौर-तरीकों में क्रॉस-भाषा अनुवाद और ज़ीरो-शॉट वर्गीकरण के लिए टूल प्रदान करता है। फ्रेमवर्क एंटिटी पहचान और भावना विश्लेषण के लिए पर्यवेक्षित मॉडल प्रशिक्षण, साथ ही निष्कर्षण प्रश्न उत्तर और दस्तावेज़ सारांश सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह समानता खोज के लिए वेक्टर डेटाबेस समर्थन को एकीकृत करता है और GPU त्वरण और केंद्रीकृत रजिस्ट्री के माध्यम से मॉडल लाइफसाइकिल प्रबंधन के लिए बुनियादी ढांचा प्रदान करता है। टूलकिट एक सार्वजनिक रिपॉजिटरी के माध्यम से कस्टम मॉडल और पाइपलाइनों के वितरण की अनुमति देता है और REST API के माध्यम से मॉडल की तैनाती का समर्थन करता है।

    Combines visual and textual data into a shared embedding space for image captioning and document reasoning.

    Scala
    GitHub पर देखें↗4,135
  • apple/ml-mgieapple का अवतार

    apple/ml-mgie

    3,876GitHub पर देखें↗

    ml-mgie is a multimodal machine learning framework and image editor designed for instruction-based image manipulation. It utilizes multimodal large language models to translate natural language prompts into precise visual modifications, functioning as a text-to-image editing model. The system is a research implementation focused on aligning visual imagination with textual commands. It employs a training process based on image-pair datasets and descriptive instructions to learn how to execute complex visual edits. The framework covers capabilities in AI-powered visual content creation, includ

    Integrates visual and textual encoders to interpret editing instructions and generate modification parameters.

    Python
    GitHub पर देखें↗3,876
  1. Home
  2. Artificial Intelligence & ML
  3. Visual-Language Multimodal Integration