Zonos is a controllable audio synthesis engine and large language model for text-to-speech. It serves as a multilingual speech generator capable of producing audio in English, Japanese, Chinese, French, and German. The system provides zero-shot voice cloning, allowing the replication of specific human voices using short audio samples. It supports the capture of nuanced behaviors, such as whispering, and provides parametric control over speaking rate, pitch, frequency, and emotional tone. The project covers a broad range of expressive speech synthesis and custom audio generation capabilities,
MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English, including seamless code-switching within a single utterance. It functions as a text-to-speech engine, voice cloning system, and speech-to-text alignment tool, built around an acoustic latent compression model that encodes high-resolution audio into compact representations for efficient processing. The system distinguishes itself through accent intensity control, allowing adjustment of a speaker's accent strength in generated speech, and voice cloning from short audio samples for pers
MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions as a voice synthesis trainer for building custom models from audio datasets, a command-line generator for producing audio files, and a text-to-speech server for remote application integration. The project specializes in real-time voice cloning, which extracts vocal characteristics from short audio samples to mimic a target speaker's unique timbre. It utilizes reference-driven audio synthesis to condition pre-trained models on specific audio samples, allowing for the generation
Acest proiect este un model de bază expresiv de text-to-speech și un sistem de clonare a vocii, conceput pentru a sintetiza vorbirea umană cu nuanțe emoționale și înaltă fidelitate. Acesta funcționează ca un model de vorbire finetunabil care poate genera audio imitând o anumită persoană folosind o mostră de voce de referință.
Principalele funcționalități ale metavoiceio/metavoice-src sunt: Zero-Shot Voice Cloning, Foundation Models, Expressive Synthesis, Synthesis Model Finetuning, Finetuning Workflows, Model Finetuning, Speech Synthesis Models, Expressive Speech Synthesis.
Alternativele open-source pentru metavoiceio/metavoice-src includ: zyphra/zonos — Zonos is a controllable audio synthesis engine and large language model for text-to-speech. It serves as a… getstream/vision-agents. bytedance/megatts3 — MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English,… jasonppy/voicecraft — VoiceCraft is a neural speech generation and manipulation system consisting of a text-to-speech system, a voice… babysor/mockingbird — MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions… neonbjb/tortoise-tts — Tortoise-tts is a neural text-to-speech engine and voice cloning toolkit designed for high-quality audio generation.…