awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेसMCP सर्वर
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
bytedance avatar

bytedance/MegaTTS3

0
View on GitHub↗
6,066 स्टार्स·469 फोर्क्स·Python·apache-2.0·9 व्यूज़

MegaTTS3

MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English, including seamless code-switching within a single utterance. It functions as a text-to-speech engine, voice cloning system, and speech-to-text alignment tool, built around an acoustic latent compression model that encodes high-resolution audio into compact representations for efficient processing.

The system distinguishes itself through accent intensity control, allowing adjustment of a speaker's accent strength in generated speech, and voice cloning from short audio samples for personalized synthesis. It provides both a command-line interface for automated speech generation without a graphical environment and a web-based inference UI for browser-driven voice sample upload and text-to-speech output. A pseudo-label aligner trains text-speech alignment models using expert-generated labels for robust alignment.

Additional capabilities include grapheme-to-phoneme conversion for improved pronunciation accuracy, latent diffusion transformer-based audio reconstruction, and support for bilingual speech synthesis with code-switching. The system compresses speech into acoustic latents for efficient storage and downstream voice conversion tasks.

Features

  • Bilingual Speech Synthesizers - Generates natural-sounding speech in Chinese and English, including code-switching within a single utterance.
  • Text-to-Speech Engines - Converts written text into natural-sounding speech using a lightweight diffusion transformer model.
  • Latent Space Encoders - Encodes high-quality audio into a compact latent representation that can be reconstructed with minimal loss.
  • Zero-Shot Voice Cloning - Replicates a speaker's voice using only a brief audio reference for personalized speech synthesis.
  • Grapheme To Phoneme Conversion - Converts written text into phonetic representations for improved pronunciation accuracy.
  • Acoustic-Text Alignment - Aligns spoken audio to its corresponding text using a robust aligner trained on pseudo-labels from expert models.
  • Text-to-Speech - Converts written text into natural-sounding speech using a lightweight diffusion transformer model.
  • Latent Acoustic Mapping - Encodes high-resolution audio into a compact latent representation for efficient model training and voice conversion.
  • Acoustic Latent Compressors - Encodes high-resolution audio into a compact latent representation for efficient model training and voice conversion.
  • Command-Line Speech Synthesizers - Accepts a voice sample and text as arguments to produce speech output without a graphical interface.
  • Speech Latent - Encodes speech into a compact latent space and reconstructs audio using a diffusion-based transformer decoder.
  • Voice Cloning - Replicates a speaker's voice characteristics using only a brief audio reference, enabling personalized speech synthesis.
  • Speech Accent Transformation - Provides accent intensity control by scaling learned accent embeddings during inference.
  • Bilingual Code-Switching - Supports seamless switching between Chinese and English within a single utterance.
  • Bilingual Synthesizers - Generates speech in Chinese and English, including code-switching within a single utterance.
  • Speech-Text Pseudo-Label Aligners - Aligns spoken audio with its corresponding text transcription using a robust aligner trained on pseudo-labels from expert models.
  • Command-Line Speech Synthesizers - Runs speech synthesis from a command line by providing a voice sample and text as arguments.
  • Web-Based Speech Synthesizers - Provides a browser-based UI for uploading voice samples and generating speech from text.
  • Web-Based Speech Inference UIs - Provides a browser interface for uploading voice samples and generating speech from text.
  • Accent Intensity Controllers - Adjusts the strength of a speaker's accent in the generated speech through configurable weight parameters.
  • Command Line Interfaces - Ships a command-line interface for generating speech from text and voice samples.
  • Speech Processing - High-performance text-to-speech model.
  • Speech Synthesis - High-performance speech synthesis model.

स्टार हिस्ट्री

bytedance/megatts3 के लिए स्टार हिस्ट्री चार्टbytedance/megatts3 के लिए स्टार हिस्ट्री चार्ट

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Start searching with AI

अक्सर पूछे जाने वाले प्रश्न

bytedance/megatts3 क्या करता है?

MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English, including seamless code-switching within a single utterance. It functions as a text-to-speech engine, voice cloning system, and speech-to-text alignment tool, built around an acoustic latent compression model that encodes high-resolution audio into compact representations for efficient processing.

bytedance/megatts3 की मुख्य विशेषताएं क्या हैं?

bytedance/megatts3 की मुख्य विशेषताएं हैं: Bilingual Speech Synthesizers, Text-to-Speech Engines, Latent Space Encoders, Zero-Shot Voice Cloning, Grapheme To Phoneme Conversion, Acoustic-Text Alignment, Text-to-Speech, Latent Acoustic Mapping।

bytedance/megatts3 के कुछ ओपन-सोर्स विकल्प क्या हैं?

bytedance/megatts3 के ओपन-सोर्स विकल्पों में शामिल हैं: netease-youdao/emotivoice — EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio… openbmb/voxcpm — VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice… myshell-ai/openvoice — OpenVoice is a multilingual text-to-speech framework and voice cloning AI model designed for high-fidelity voice… swivid/f5-tts — F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent… fishaudio/fish-speech — This project is a generative speech synthesis engine that converts text into high-fidelity human speech. It utilizes a… metavoiceio/metavoice-src — This project is an expressive text-to-speech foundation model and voice cloning system designed to synthesize…

MegaTTS3 के ओपन-सोर्स विकल्प

समान ओपन-सोर्स प्रोजेक्ट्स, जो MegaTTS3 के साथ साझा की गई सुविधाओं के आधार पर रैंक किए गए हैं।
  • netease-youdao/emotivoicenetease-youdao का अवतार

    netease-youdao/EmotiVoice

    8,446GitHub पर देखें↗

    EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio in English and Chinese. It utilizes a deep learning architecture to produce high-fidelity speech with controllable emotional states and timbres. The project includes a voice cloning framework for replicating specific speaker identities by training custom acoustic models on personal audio datasets. It employs a jointly-trained acoustic-vocoder pipeline and style-embedding-based synthesis to manage expression and reduce audio artifacts. The system covers a broad range of speec

    Pythonaideep-learningemotion
    GitHub पर देखें↗8,446
  • openbmb/voxcpmOpenBMB का अवतार

    OpenBMB/VoxCPM

    29,985GitHub पर देखें↗

    VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice cloning tool and a synthetic voice designer, capable of generating natural speech across global languages and regional dialects using a GPU-accelerated audio generator. The project features a speech model fine-tuning framework that supports both full parameter updates and low-rank adaptation for customizing voice characteristics. It enables high-fidelity voice cloning from reference audio, including cross-lingual voice transfer and acoustic environment mimicry, as well as the crea

    Pythonaudiodeeplearningminicpm
    GitHub पर देखें↗29,985
  • myshell-ai/openvoicemyshell-ai का अवतार

    myshell-ai/OpenVoice

    36,720GitHub पर देखें↗

    OpenVoice is a multilingual text-to-speech framework and voice cloning AI model designed for high-fidelity voice replication and low-latency audio generation. It functions as an instant speech synthesis engine that converts text to audio while replicating a specific speaker's tone and color. The system is distinguished by its ability to perform cross-lingual cloning, allowing the vocal characteristics of a reference speaker to be applied to speech in different languages regardless of the original training data. It utilizes a decoupled representation to separate the physical identity of a voic

    Pythontext-to-speechttsvoice-clone
    GitHub पर देखें↗36,720
  • swivid/f5-ttsSWivid का अवतार

    SWivid/F5-TTS

    14,798GitHub पर देखें↗

    F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent synthetic speech. It functions as a multilingual speech synthesizer and neural training framework, providing tools for voice cloning and high-performance inference serving. The project distinguishes itself through a voice cloning toolkit capable of mimicking specific speaker characteristics and tones from reference audio clips. It supports cross-lingual generation, allowing for the synthesis of audio across various global languages or the mixing of multiple languages within a s

    Python
    GitHub पर देखें↗14,798
  • MegaTTS3 के सभी 30 विकल्प देखें→