awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

12 रिपॉजिटरी

Awesome GitHub RepositoriesSequence Alignment Models

Attention-based architectures for dynamic alignment in sequence-to-sequence tasks.

Distinguishing note: Focuses on specific attention-based alignment techniques like Bahdanau.

Explore 12 awesome GitHub repositories matching artificial intelligence & ml · Sequence Alignment Models. Refine with filters or upvote what's useful.

Awesome Sequence Alignment Models GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • facebookresearch/fairseqfacebookresearch का अवतार

    facebookresearch/fairseq

    32,228GitHub पर देखें↗

    Fairseq is a PyTorch toolkit for sequence-to-sequence modeling, specializing in neural machine translation, automatic speech recognition, and large-scale language model training. It provides a framework for processing and aligning diverse data sources, including text, audio, and video, to support tasks such as speech-to-text conversion and multimodal sequence learning. The project is distinguished by its distributed training capabilities, which utilize parameter sharding, mixed-precision training, and CPU offloading to handle models that exceed single-device memory. It also includes specializ

    Generates mappings between source and target tokens using attention-based alignment architectures.

    Python
    GitHub पर देखें↗32,228
  • funaudiollm/cosyvoiceFunAudioLLM का अवतार

    FunAudioLLM/CosyVoice

    21,673GitHub पर देखें↗

    CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus

    Maps text inputs to specific phonetic sequences to ensure precise pronunciation and prosodic rendering.

    Pythonaudio-generationcantonesechatbot
    GitHub पर देखें↗21,673
  • m-bain/whisperxm-bain का अवतार

    m-bain/whisperX

    20,228GitHub पर देखें↗

    WhisperX is an automated speech recognition toolkit designed to convert spoken audio into text while maintaining precise synchronization with the original media. It functions as an integrated pipeline that combines transcription, phoneme-based alignment, and speaker diarization to produce structured, attributed transcripts. The project distinguishes itself through its use of forced alignment, which matches existing text to audio signals at the phoneme level to generate accurate word-level timestamps. It also incorporates speaker diarization to identify and label unique voices within a recordi

    Improves transcription accuracy by matching text to audio signals at the phoneme level.

    Pythonasrspeechspeech-recognition
    GitHub पर देखें↗20,228
  • index-tts/index-ttsindex-tts का अवतार

    index-tts/index-tts

    18,851GitHub पर देखें↗

    Index-tts is a neural audio generation engine designed to convert written text into high-fidelity human speech. By utilizing deep learning models and phoneme-based sequence modeling, the system transforms text into natural-sounding audio waveforms suitable for a variety of accessibility and media applications. The platform functions as a server-side inference pipeline that provides a programmatic interface for integrating voice generation into external applications. It distinguishes itself through asynchronous audio streaming, which buffers and delivers generated speech chunks in real time to

    Converts raw text into structured phonetic units to ensure accurate pronunciation and natural prosody.

    Pythonbigvgancross-lingualindextts
    GitHub पर देखें↗18,851
  • rhasspy/piperrhasspy का अवतार

    rhasspy/piper

    10,584GitHub पर देखें↗

    Piper is a local neural text-to-speech engine designed to convert written text into natural human speech entirely on your own hardware. By utilizing a neural synthesis framework, it operates without the need for internet connectivity, ensuring that all audio generation remains private and secure. The system distinguishes itself through a modular architecture that allows for the dynamic loading of speaker embeddings and voice configurations. This enables users to switch between various vocal personas and styles without requiring a full reload of the core synthesis model. By processing input th

    Processes input through a phoneme-based pipeline to ensure consistent pronunciation and accurate prosody.

    C++speech-synthesistext-to-speechtts
    GitHub पर देखें↗10,584
  • jasonppy/voicecraftjasonppy का अवतार

    jasonppy/VoiceCraft

    8,500GitHub पर देखें↗

    VoiceCraft is a neural speech generation and manipulation system consisting of a text-to-speech system, a voice cloning tool, and an audio inpainting engine. It uses a large language model approach to synthesize high-fidelity audio from text and replicate speaker identities. The system provides zero-shot voice cloning and speech editing capabilities, allowing users to modify spoken content within existing recordings. This includes an audio inpainting engine that replaces specific sections of audio with new speech while preserving the original acoustic characteristics and speaker identity. Th

    Converts text and audio transcripts into discrete phonetic units to standardize speech generation.

    Jupyter Notebook
    GitHub पर देखें↗8,500
  • netease-youdao/emotivoicenetease-youdao का अवतार

    netease-youdao/EmotiVoice

    8,446GitHub पर देखें↗

    EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio in English and Chinese. It utilizes a deep learning architecture to produce high-fidelity speech with controllable emotional states and timbres. The project includes a voice cloning framework for replicating specific speaker identities by training custom acoustic models on personal audio datasets. It employs a jointly-trained acoustic-vocoder pipeline and style-embedding-based synthesis to manage expression and reduce audio artifacts. The system covers a broad range of speec

    Implements a pipeline to transform raw bilingual text into phonetic representations for synthesis.

    Pythonaideep-learningemotion
    GitHub पर देखें↗8,446
  • multimodal-art-projection/yuemultimodal-art-projection का अवतार

    multimodal-art-projection/YuE

    6,292GitHub पर देखें↗

    YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open

    Aligns phoneme-level lyric timing with generated musical notes using a cross-attention mechanism between text and audio tokens.

    Pythonaiaudio-generationdeep-learning
    GitHub पर देखें↗6,292
  • zhaochenyang20/awesome-ml-sys-tutorialzhaochenyang20 का अवतार

    zhaochenyang20/Awesome-ML-SYS-Tutorial

    5,371GitHub पर देखें↗

    This project provides a comprehensive technical guide and framework for engineering large-scale machine learning systems. It covers the full lifecycle of model development, focusing on the infrastructure and computational principles required to build, train, and serve generative AI models across distributed GPU clusters. The repository distinguishes itself by offering deep-dive tutorials and implementation strategies for complex system challenges. It emphasizes high-performance architectural primitives, such as collective communication orchestration, distributed tensor sharding, and static gr

    Implements attention-based architectures for dynamic alignment in sequence-to-sequence tasks.

    Python
    GitHub पर देखें↗5,371
  • microsoft/muzicmicrosoft का अवतार

    microsoft/muzic

    4,928GitHub पर देखें↗

    Muzic AI-संचालित संगीत विश्लेषण, रचना और संश्लेषण के लिए एक डीप लर्निंग प्लेटफ़ॉर्म और फ्रेमवर्क है। यह एक संगीत जनरेशन फ्रेमवर्क और विश्लेषण टूल के रूप में कार्य करता है, जो प्रतीकात्मक और ऑडियो संगीत के निर्माण और व्याख्या को व्यवस्थित करने के लिए बड़े भाषा मॉडल्स और स्वायत्त एजेंटों का उपयोग करता है। यह प्रोजेक्ट अपनी क्रॉस-मॉडल क्षमताओं द्वारा प्रतिष्ठित है, जो ज़ीरो-शॉट वर्गीकरण और सूचना पुनर्प्राप्ति के लिए प्राकृतिक भाषा और प्रतीकात्मक संगीत को एक साझा संयुक्त एम्बेडिंग स्पेस में मैप करता है। यह विभिन्न प्रकार के विशेष आर्किटेक्चर को नियोजित करता है, जिसमें ऑडियो संश्लेषण के लिए डिफ्यूज़न फ्रेमवर्क, लंबी-अनुक्रम संरचनात्मक स्थिरता के लिए डुअल-ग्रेन अटेंशन मैकेनिज्म और एक हाइब्रिड सिस्टम शामिल है जो न्यूरल नेटवर्क के साथ संगीत सिद्धांत नियमों को जोड़ता है। यह प्लेटफ़ॉर्म टेक्स्ट और लिरिक्स से MIDI अनुक्रमों के निर्माण, न्यूरल सिंगिंग वॉयस सिंथेसिस और स्वचालित लिरिक्स ट्रांसक्रिप्शन सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह संगीत संरचना मॉडलिंग, विशेषता-आधारित प्रतीकात्मक जनरेशन और स्वायत्त एजेंटों के माध्यम से बाहरी संगीत टूल्स के ऑर्केस्ट्रेशन के लिए टूल्स भी प्रदान करता है। सहायक यूटिलिटीज में बड़े पैमाने पर MIDI बाइनराइजेशन, डेटासेट एन्कोडिंग और मेलोडी नोट निष्कर्षण और स्पीच-टू-फोनम एलाइनमेंट के लिए ऑडियो सिग्नल प्रोसेसिंग के लिए डेटा इंजीनियरिंग पाइपलाइन शामिल हैं।

    Determines the exact timing of phonemes within a speech audio signal to facilitate syllable-level adjustments.

    Pythonai-musicdeep-learningmusic
    GitHub पर देखें↗4,928
  • andabi/deep-voice-conversionandabi का अवतार

    andabi/deep-voice-conversion

    3,941GitHub पर देखें↗

    यह प्रोजेक्ट TensorFlow पर आधारित एक वॉइस कन्वर्जन फ्रेमवर्क और डीप लर्निंग ऑडियो टूलकिट है, जिसे न्यूरल वॉइस स्टाइल ट्रांसफर के लिए बनाया गया है। यह एक स्पीच सिंथेसिस इंजन के रूप में काम करता है जो सोर्स स्पीकर की आवाज़ की स्पेक्ट्रल विशेषताओं को टारगेट स्पीकर की आवाज़ में बदल देता है। यह सिस्टम वॉइस कन्वर्जन के लिए फोनम-आधारित (phoneme-based) दृष्टिकोण अपनाता है, जो ऑडियो को स्पीकर-इंडिपेंडेंट फोनम में वर्गीकृत करता है और फिर उन्हें टारगेट वॉइस का उपयोग करके फिर से सिंथेसाइज करता है। यह पाइपलाइन अलग-अलग स्पीकर्स के बीच ऑडियो फीचर्स को मैप करके वॉइस विशेषताओं को बदलने की सुविधा देती है। इस टूलकिट में मल्टीपल GPUs पर ऑडियो मॉडल ट्रेनिंग, टेंसर डेटा नॉर्मलाइजेशन और मॉडल हाइपरपैरामीटर्स के प्रबंधन की क्षमताएं शामिल हैं। यह परफॉरमेंस मॉनिटरिंग के लिए भी टूल्स प्रदान करता है, जैसे कि कन्फ्यूजन मैट्रिक्स के जरिए क्लासिफिकेशन एक्यूरेसी को विज़ुअलाइज़ करना।

    Transforms audio by analyzing speaker-independent phonemes and resynthesizing them using a target voice.

    Python
    GitHub पर देखें↗3,941
  • voicevox/voicevoxVOICEVOX का अवतार

    VOICEVOX/voicevox

    3,025GitHub पर देखें↗

    Voicevox is a text-to-speech synthesis software and audio production environment that converts written text into spoken audio using synthetic character voices. It functions as both a comprehensive editor for voice design and a standalone speech synthesis engine capable of generating audio via an API for integration into external applications. The project distinguishes itself by providing a singing voice synthesizer that uses a piano-roll interface for melodic vocal composition, including the ability to generate humming. It offers specialized prosody editing tools for the manual refinement of

    Uses a customizable dictionary-based system to translate written text into phonetic representations for accurate pronunciation.

    TypeScript
    GitHub पर देखें↗3,025
  1. Home
  2. Artificial Intelligence & ML
  3. Sequence Alignment Models

सब-टैग एक्सप्लोर करें

  • Phoneme-Based Alignment3 सब-टैग्सAlignment models that use phoneme-level analysis to match text to audio. **Distinct from Sequence Alignment Models:** Distinct from general sequence alignment: focuses on phoneme-level acoustic matching for transcription accuracy.