1 مستودع
Encodes high-resolution audio into compact latent representations for efficient model training and voice conversion.
Distinct from Latent Acoustic Mapping: Distinct from general Latent Acoustic Mapping: focuses on compression of audio into latents, not mapping text to latent space.
Explore 1 awesome GitHub repository matching artificial intelligence & ml · Acoustic Latent Compressors. Refine with filters or upvote what's useful.
MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English, including seamless code-switching within a single utterance. It functions as a text-to-speech engine, voice cloning system, and speech-to-text alignment tool, built around an acoustic latent compression model that encodes high-resolution audio into compact representations for efficient processing. The system distinguishes itself through accent intensity control, allowing adjustment of a speaker's accent strength in generated speech, and voice cloning from short audio samples for pers
Encodes high-resolution audio into a compact latent representation for efficient model training and voice conversion.