13 dépôts
Models capable of synthesizing speech across multiple languages within a single system.
Distinct from Speech Synthesis Models: Focuses on the multilingual capability of synthesis, whereas Speech Synthesis Models is the general architecture.
Explore 13 awesome GitHub repositories matching artificial intelligence & ml · Multilingual Synthesis. Refine with filters or upvote what's useful.
ChatTTS is a conversational text-to-speech generative model designed to convert written dialogue into natural sounding audio. It functions as a multilingual speech synthesis framework capable of producing human-like audio across different languages and speaker profiles. The system is distinguished by its ability to generate interactive dialogue with realistic vocal nuances. It utilizes a speech nuance controller to insert specific tokens that trigger non-verbal elements, such as laughter, pauses, and interjections, during the synthesis process. The project includes a streaming audio generato
Provides a generative framework capable of producing human-like speech across multiple languages.
OpenVoice is a multilingual text-to-speech framework and voice cloning AI model designed for high-fidelity voice replication and low-latency audio generation. It functions as an instant speech synthesis engine that converts text to audio while replicating a specific speaker's tone and color. The system is distinguished by its ability to perform cross-lingual cloning, allowing the vocal characteristics of a reference speaker to be applied to speech in different languages regardless of the original training data. It utilizes a decoupled representation to separate the physical identity of a voic
Synthesizes natural sounding speech across multiple languages while preserving a specific speaker's unique characteristics.
Readest is a cross-platform digital book reader and library management system designed to render multi-format ebook files across different devices. It provides a consistent interface for viewing digital content while coordinating reading progress, bookmarks, and notes through synchronization services. The application includes specialized tools for technical and academic reading, such as code syntax highlighting, a virtual split-screen viewport for comparing documents, and full-text search for rapid information retrieval. It further extends the reading experience with integrated dictionary loo
Generates multilingual narration of written content using integrated voice synthesis technology.
F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent synthetic speech. It functions as a multilingual speech synthesizer and neural training framework, providing tools for voice cloning and high-performance inference serving. The project distinguishes itself through a voice cloning toolkit capable of mimicking specific speaker characteristics and tones from reference audio clips. It supports cross-lingual generation, allowing for the synthesis of audio across various global languages or the mixing of multiple languages within a s
Synthesizes spoken audio across global languages using specialized model checkpoints for different regions.
Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device voice synthesis tool, and a speaker identification framework. It is designed as a cross-platform speech API that enables speech-to-text, text-to-speech, and speaker verification tasks to be executed locally on a device without requiring network access. The project is distinguished by its ability to perform zero-shot voice cloning and speaker diarization on-device. It supports a wide range of hardware accelerations, including GPU and various NPU architectures, and provides a Web
Converts mixed language text into audio using a multi-speaker model.
This project is a framework for developing multimodal AI agents that function as programmable participants in real-time communication rooms. It enables the construction of agents that can see, hear, and speak by integrating speech-to-text, large language models, and text-to-speech pipelines to facilitate low-latency, natural conversations. The system is distinguished by its advanced orchestration of real-time media and conversational flow, including support for full-duplex speech, preemptive response generation, and sophisticated interruption management. It further differentiates itself throu
Synthesizes natural-sounding audio across various regional and English languages with configurable styles.
Bert-VITS2 is a neural speech synthesis system and AI voice generator designed to convert written text into natural sounding audio. It utilizes a VITS2 engine and a neural speech synthesis model to produce high-fidelity human voices. The system incorporates a multilingual BERT language processor to improve the prosody and emotional accuracy of the generated speech. It supports multilingual voice generation and custom voice cloning to replicate specific human speech patterns and tones. The architecture covers text-to-speech synthesis through a multi-stage pipeline involving phoneme alignment,
Features a multilingual synthesis architecture capable of generating spoken audio in multiple different languages.
VALL-E-X is a neural speech synthesis framework and zero-shot text-to-speech engine. It functions as a multilingual synthesizer capable of generating natural human speech with control over emotion, pitch, and prosody. The project specializes in zero-shot voice cloning and cross-lingual voice replication, allowing the system to produce personalized speech in multiple target languages using short audio samples without additional training. It further enables cross-language accent manipulation and the ability to match the emotional tone and acoustic environment of a provided prompt. The implemen
Provides a framework capable of synthesizing expressive audio across multiple languages within a single system.
Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large language model architectures. It functions as a multilingual speech synthesizer capable of generating high-fidelity audio across different languages with control over emotional tone and prosody. The system includes a voice cloning tool that creates synthetic replicas of specific speakers from short audio samples without requiring extensive model training. It also provides a streaming audio API designed to deliver generated speech incrementally to minimize playback delay. The
Synthesizes natural-sounding spoken audio across multiple languages within a single generative system.
ChatTTS-ui est une interface web et un wrapper d'API pour le modèle ChatTTS, conçu pour convertir du texte écrit et des entrées multilingues en audio parlé. Il fonctionne comme un tableau de bord de synthèse vocale par IA et un générateur programmatique pour créer une sortie vocale naturaliste. Le projet se concentre sur le profilage vocal personnalisé et le contrôle des nuances de la parole. Il permet de maintenir des caractéristiques de locuteur cohérentes en utilisant des valeurs de seed et des fichiers de données, tout en offrant des contrôles pour le ton, le rire et les pauses via des prompts comportementaux et des paramètres d'échantillonnage. Le système inclut une architecture client-serveur qui gère le traitement audio asynchrone et fournit une interface programmatique pour l'intégration d'applications externes. Il gère les profils vocaux et les configurations audio via une interface à état géré pour assurer une synthèse cohérente.
Transforms mixed language text into spoken audio through a unified synthesis interface.
Il s'agit d'une collection de modèles neuronaux pré-entraînés pour la reconnaissance vocale, la synthèse et la détection d'activité vocale. Elle fournit une bibliothèque d'actifs conçus pour la conversion parole-texte, la synthèse texte-parole et l'identification de segments de parole humaine au sein de l'audio. Le projet propose une synthèse texte-parole avec prise en charge de plusieurs langues et l'utilisation du langage de balisage de synthèse vocale (SSML) pour contrôler la prosodie, la hauteur et le timing. Pour la reconnaissance vocale, le système inclut des capacités de transcription audio en texte avec extraction d'horodatage au niveau du mot et un restaurateur de ponctuation automatisé pour insérer les majuscules et la ponctuation dans le texte brut. Les modèles sont exportés au format Open Neural Network Exchange et TorchScript pour permettre une exécution haute performance sur différents accélérateurs matériels et systèmes d'exploitation.
Provides synthesis models capable of generating speech across a wide variety of regional and minority languages.
Bark Voice Cloning est un moteur de synthèse texte-parole conçu pour générer un audio au son naturel et répliquer des caractéristiques vocales spécifiques. Le système utilise un modèle autorégressif basé sur transformer pour convertir le texte écrit en parole haute fidélité, prenant en charge la sortie multilingue et une livraison expressive. Le projet se distingue par le clonage de voix zero-shot, qui extrait les embeddings d'identité du locuteur à partir de courts échantillons audio pour conditionner le modèle génératif sans nécessiter de réglage fin étendu. Il fournit également des flux de travail spécialisés pour la conversion d'identité vocale, permettant aux utilisateurs de transformer le locuteur d'un enregistrement existant tout en préservant la livraison émotionnelle et les modèles rythmiques originaux. La plateforme englobe une suite complète d'outils pour la synthèse vocale et la manipulation audio. Cela inclut des utilitaires pour extraire l'audio source des médias, entraîner des modèles vocaux personnalisés et mapper le contenu linguistique sémantique vers des jetons acoustiques fins. Le logiciel est distribué sous forme d'une collection de notebooks Jupyter qui facilitent l'exécution de ces pipelines d'inférence multi-étapes.
Generates natural-sounding audio in multiple languages from text input using specialized phonetic models.
This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of conversational AI agents. It enables the creation of lifelike spoken audio from text, the replication of human voices through custom cloning, and the deployment of real-time voice agents capable of interacting with external large language models. The library distinguishes itself through deep integration of conversational AI capabilities, including the design of agent personas and the execution of real-time actions via APIs. It supports professional-grade audio production thro
Provides high-fidelity speech synthesis across multiple languages while maintaining native-level emotion and clarity.