awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
bytedance avatar

bytedance/MegaTTS3

0
View on GitHub↗
6,066 نجوم·469 تفرعات·Python·apache-2.0·13 مشاهدات

MegaTTS3

MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English, including seamless code-switching within a single utterance. It functions as a text-to-speech engine, voice cloning system, and speech-to-text alignment tool, built around an acoustic latent compression model that encodes high-resolution audio into compact representations for efficient processing.

The system distinguishes itself through accent intensity control, allowing adjustment of a speaker's accent strength in generated speech, and voice cloning from short audio samples for personalized synthesis. It provides both a command-line interface for automated speech generation without a graphical environment and a web-based inference UI for browser-driven voice sample upload and text-to-speech output. A pseudo-label aligner trains text-speech alignment models using expert-generated labels for robust alignment.

Additional capabilities include grapheme-to-phoneme conversion for improved pronunciation accuracy, latent diffusion transformer-based audio reconstruction, and support for bilingual speech synthesis with code-switching. The system compresses speech into acoustic latents for efficient storage and downstream voice conversion tasks.

Features

  • Bilingual Speech Synthesizers - Generates natural-sounding speech in Chinese and English, including code-switching within a single utterance.
  • Text-to-Speech Engines - Converts written text into natural-sounding speech using a lightweight diffusion transformer model.
  • Latent Space Encoders - Encodes high-quality audio into a compact latent representation that can be reconstructed with minimal loss.
  • Zero-Shot Voice Cloning - Replicates a speaker's voice using only a brief audio reference for personalized speech synthesis.
  • Grapheme To Phoneme Conversion - Converts written text into phonetic representations for improved pronunciation accuracy.
  • Acoustic-Text Alignment - Aligns spoken audio to its corresponding text using a robust aligner trained on pseudo-labels from expert models.
  • Text-to-Speech - Converts written text into natural-sounding speech using a lightweight diffusion transformer model.
  • Latent Acoustic Mapping - Encodes high-resolution audio into a compact latent representation for efficient model training and voice conversion.
  • Acoustic Latent Compressors - Encodes high-resolution audio into a compact latent representation for efficient model training and voice conversion.
  • Command-Line Speech Synthesizers - Accepts a voice sample and text as arguments to produce speech output without a graphical interface.
  • Speech Latent - Encodes speech into a compact latent space and reconstructs audio using a diffusion-based transformer decoder.
  • Voice Cloning - Replicates a speaker's voice characteristics using only a brief audio reference, enabling personalized speech synthesis.
  • Speech Accent Transformation - Provides accent intensity control by scaling learned accent embeddings during inference.
  • Bilingual Code-Switching - Supports seamless switching between Chinese and English within a single utterance.
  • Bilingual Synthesizers - Generates speech in Chinese and English, including code-switching within a single utterance.
  • Speech-Text Pseudo-Label Aligners - Aligns spoken audio with its corresponding text transcription using a robust aligner trained on pseudo-labels from expert models.
  • Command-Line Speech Synthesizers - Runs speech synthesis from a command line by providing a voice sample and text as arguments.
  • Web-Based Speech Synthesizers - Provides a browser-based UI for uploading voice samples and generating speech from text.
  • Web-Based Speech Inference UIs - Provides a browser interface for uploading voice samples and generating speech from text.
  • Accent Intensity Controllers - Adjusts the strength of a speaker's accent in the generated speech through configurable weight parameters.
  • Command Line Interfaces - Ships a command-line interface for generating speech from text and voice samples.
  • Speech Processing - High-performance text-to-speech model.
  • Speech Synthesis - High-performance speech synthesis model.

سجل النجوم

مخطط تاريخ النجوم لـ bytedance/megatts3مخطط تاريخ النجوم لـ bytedance/megatts3

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

بدائل مفتوحة المصدر لـ MegaTTS3

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع MegaTTS3.
  • netease-youdao/emotivoiceالصورة الرمزية لـ netease-youdao

    netease-youdao/EmotiVoice

    8,446عرض على GitHub↗

    EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio in English and Chinese. It utilizes a deep learning architecture to produce high-fidelity speech with controllable emotional states and timbres. The project includes a voice cloning framework for replicating specific speaker identities by training custom acoustic models on personal audio datasets. It employs a jointly-trained acoustic-vocoder pipeline and style-embedding-based synthesis to manage expression and reduce audio artifacts. The system covers a broad range of speec

    Pythonaideep-learningemotion
    عرض على GitHub↗8,446
  • openbmb/voxcpmالصورة الرمزية لـ OpenBMB

    OpenBMB/VoxCPM

    29,985عرض على GitHub↗

    VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice cloning tool and a synthetic voice designer, capable of generating natural speech across global languages and regional dialects using a GPU-accelerated audio generator. The project features a speech model fine-tuning framework that supports both full parameter updates and low-rank adaptation for customizing voice characteristics. It enables high-fidelity voice cloning from reference audio, including cross-lingual voice transfer and acoustic environment mimicry, as well as the crea

    Pythonaudiodeeplearningminicpm
    عرض على GitHub↗29,985
  • myshell-ai/openvoiceالصورة الرمزية لـ myshell-ai

    myshell-ai/OpenVoice

    36,720عرض على GitHub↗

    OpenVoice is a multilingual text-to-speech framework and voice cloning AI model designed for high-fidelity voice replication and low-latency audio generation. It functions as an instant speech synthesis engine that converts text to audio while replicating a specific speaker's tone and color. The system is distinguished by its ability to perform cross-lingual cloning, allowing the vocal characteristics of a reference speaker to be applied to speech in different languages regardless of the original training data. It utilizes a decoupled representation to separate the physical identity of a voic

    Pythontext-to-speechttsvoice-clone
    عرض على GitHub↗36,720
  • swivid/f5-ttsالصورة الرمزية لـ SWivid

    SWivid/F5-TTS

    14,798عرض على GitHub↗

    F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent synthetic speech. It functions as a multilingual speech synthesizer and neural training framework, providing tools for voice cloning and high-performance inference serving. The project distinguishes itself through a voice cloning toolkit capable of mimicking specific speaker characteristics and tones from reference audio clips. It supports cross-lingual generation, allowing for the synthesis of audio across various global languages or the mixing of multiple languages within a s

    Python
    عرض على GitHub↗14,798
عرض جميع البدائل الـ 30 لـ MegaTTS3→

الأسئلة الشائعة

ما هي وظيفة bytedance/megatts3؟

MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English, including seamless code-switching within a single utterance. It functions as a text-to-speech engine, voice cloning system, and speech-to-text alignment tool, built around an acoustic latent compression model that encodes high-resolution audio into compact representations for efficient processing.

ما هي الميزات الرئيسية لـ bytedance/megatts3؟

الميزات الرئيسية لـ bytedance/megatts3 هي: Bilingual Speech Synthesizers, Text-to-Speech Engines, Latent Space Encoders, Zero-Shot Voice Cloning, Grapheme To Phoneme Conversion, Acoustic-Text Alignment, Text-to-Speech, Latent Acoustic Mapping.

ما هي البدائل مفتوحة المصدر لـ bytedance/megatts3؟

تشمل البدائل مفتوحة المصدر لـ bytedance/megatts3: netease-youdao/emotivoice — EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio… openbmb/voxcpm — VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice… myshell-ai/openvoice — OpenVoice is a multilingual text-to-speech framework and voice cloning AI model designed for high-fidelity voice… swivid/f5-tts — F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent… fishaudio/fish-speech — This project is a generative speech synthesis engine that converts text into high-fidelity human speech. It utilizes a… metavoiceio/metavoice-src — This project is an expressive text-to-speech foundation model and voice cloning system designed to synthesize…