awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

2 रिपॉजिटरी

Awesome GitHub RepositoriesAudio-Visual Semantic Alignment

Matching sounds to corresponding images or videos by mapping both to a shared mathematical space.

Distinct from Audio-Visual Signal Alignment: Distinct from temporal signal alignment as it focuses on semantic meaning rather than clock synchronization

Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Audio-Visual Semantic Alignment. Refine with filters or upvote what's useful.

Awesome Audio-Visual Semantic Alignment GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • facebookresearch/imagebindfacebookresearch का अवतार

    facebookresearch/ImageBind

    9,036GitHub पर देखें↗

    ImageBind is a multi-modal embedding model and joint representation learner that maps images, text, audio, and other modalities into a single shared vector space. It functions as a cross-modal retrieval framework designed to bind multiple sensory inputs into one cohesive mathematical embedding. The system uses a contrastive learning architecture to align disparate data types by maximizing the similarity between related samples. This allows the model to perform zero-shot multimodal classification and execute cross-modal data retrieval, such as locating visual content via natural language descr

    Matches specific sounds to corresponding images by mapping both to a common semantic space.

    Python
    GitHub पर देखें↗9,036
  • bytedance/latentsyncbytedance का अवतार

    bytedance/LatentSync

    5,806GitHub पर देखें↗

    LatentSync एक ऑडियो-ड्रिवन वीडियो जनरेटर और लेटेंट डिफ़्यूज़न लिप सिंक मॉडल है जिसे वीडियो में स्पीकर के होंठों की गतिविधियों को टारगेट ऑडियो ट्रैक के साथ सिंक्रोनाइज़ करने के लिए डिज़ाइन किया गया है। यह कस्टम वीडियो और ऑडियो डेटासेट पर सिंक्रोनाइज़ेशन नेटवर्क विकसित करने के लिए एक लिप सिंक्रोनाइज़ेशन ट्रेनिंग फ़्रेमवर्क प्रदान करता है। यह सिस्टम फेस डेटा को साफ़ करने, सेगमेंट करने और संरेखित करने के लिए एक वीडियो प्रीप्रोसेसिंग पाइपलाइन का उपयोग करता है। इसमें एक विज़ुअल सिंक मूल्यांकन टूल शामिल है जो जेनरेट किए गए वीडियो में ऑडियो और विज़ुअल संरेखण की सटीकता को मापने के लिए कॉन्फ़िडेंस स्कोर की गणना करता है। यह प्रोजेक्ट कस्टम सिंक्रोनाइज़ेशन नेटवर्क विकास, हार्डवेयर मेमोरी और रिज़ॉल्यूशन के लिए ट्रेनिंग कॉन्फ़िगरेशन मैनेजमेंट और सिंथेटिक वीडियो मूल्यांकन के लिए क्षमताओं को कवर करता है।

    Processes video and audio data to ensure facial movements match the timing and patterns of speech.

    Python
    GitHub पर देखें↗5,806
  1. Home
  2. Artificial Intelligence & ML
  3. Audio-Visual Semantic Alignment