13 مستودعات
Models capable of synthesizing speech across multiple languages within a single system.
Distinct from Speech Synthesis Models: Focuses on the multilingual capability of synthesis, whereas Speech Synthesis Models is the general architecture.
Explore 13 awesome GitHub repositories matching artificial intelligence & ml · Multilingual Synthesis. Refine with filters or upvote what's useful.
ChatTTS is a conversational text-to-speech generative model designed to convert written dialogue into natural sounding audio. It functions as a multilingual speech synthesis framework capable of producing human-like audio across different languages and speaker profiles. The system is distinguished by its ability to generate interactive dialogue with realistic vocal nuances. It utilizes a speech nuance controller to insert specific tokens that trigger non-verbal elements, such as laughter, pauses, and interjections, during the synthesis process. The project includes a streaming audio generato
Provides a generative framework capable of producing human-like speech across multiple languages.
OpenVoice is a multilingual text-to-speech framework and voice cloning AI model designed for high-fidelity voice replication and low-latency audio generation. It functions as an instant speech synthesis engine that converts text to audio while replicating a specific speaker's tone and color. The system is distinguished by its ability to perform cross-lingual cloning, allowing the vocal characteristics of a reference speaker to be applied to speech in different languages regardless of the original training data. It utilizes a decoupled representation to separate the physical identity of a voic
Synthesizes natural sounding speech across multiple languages while preserving a specific speaker's unique characteristics.
Readest is a cross-platform digital book reader and library management system designed to render multi-format ebook files across different devices. It provides a consistent interface for viewing digital content while coordinating reading progress, bookmarks, and notes through synchronization services. The application includes specialized tools for technical and academic reading, such as code syntax highlighting, a virtual split-screen viewport for comparing documents, and full-text search for rapid information retrieval. It further extends the reading experience with integrated dictionary loo
Generates multilingual narration of written content using integrated voice synthesis technology.
F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent synthetic speech. It functions as a multilingual speech synthesizer and neural training framework, providing tools for voice cloning and high-performance inference serving. The project distinguishes itself through a voice cloning toolkit capable of mimicking specific speaker characteristics and tones from reference audio clips. It supports cross-lingual generation, allowing for the synthesis of audio across various global languages or the mixing of multiple languages within a s
Synthesizes spoken audio across global languages using specialized model checkpoints for different regions.
Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device voice synthesis tool, and a speaker identification framework. It is designed as a cross-platform speech API that enables speech-to-text, text-to-speech, and speaker verification tasks to be executed locally on a device without requiring network access. The project is distinguished by its ability to perform zero-shot voice cloning and speaker diarization on-device. It supports a wide range of hardware accelerations, including GPU and various NPU architectures, and provides a Web
Converts mixed language text into audio using a multi-speaker model.
This project is a framework for developing multimodal AI agents that function as programmable participants in real-time communication rooms. It enables the construction of agents that can see, hear, and speak by integrating speech-to-text, large language models, and text-to-speech pipelines to facilitate low-latency, natural conversations. The system is distinguished by its advanced orchestration of real-time media and conversational flow, including support for full-duplex speech, preemptive response generation, and sophisticated interruption management. It further differentiates itself throu
Synthesizes natural-sounding audio across various regional and English languages with configurable styles.
Bert-VITS2 is a neural speech synthesis system and AI voice generator designed to convert written text into natural sounding audio. It utilizes a VITS2 engine and a neural speech synthesis model to produce high-fidelity human voices. The system incorporates a multilingual BERT language processor to improve the prosody and emotional accuracy of the generated speech. It supports multilingual voice generation and custom voice cloning to replicate specific human speech patterns and tones. The architecture covers text-to-speech synthesis through a multi-stage pipeline involving phoneme alignment,
Features a multilingual synthesis architecture capable of generating spoken audio in multiple different languages.
VALL-E-X is a neural speech synthesis framework and zero-shot text-to-speech engine. It functions as a multilingual synthesizer capable of generating natural human speech with control over emotion, pitch, and prosody. The project specializes in zero-shot voice cloning and cross-lingual voice replication, allowing the system to produce personalized speech in multiple target languages using short audio samples without additional training. It further enables cross-language accent manipulation and the ability to match the emotional tone and acoustic environment of a provided prompt. The implemen
Provides a framework capable of synthesizing expressive audio across multiple languages within a single system.
Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large language model architectures. It functions as a multilingual speech synthesizer capable of generating high-fidelity audio across different languages with control over emotional tone and prosody. The system includes a voice cloning tool that creates synthetic replicas of specific speakers from short audio samples without requiring extensive model training. It also provides a streaming audio API designed to deliver generated speech incrementally to minimize playback delay. The
Synthesizes natural-sounding spoken audio across multiple languages within a single generative system.
ChatTTS-ui هو واجهة ويب وغلاف لواجهة برمجة التطبيقات (API) لنموذج ChatTTS، مصمم لتحويل النصوص المكتوبة والمدخلات متعددة اللغات إلى صوت مسموع. يعمل كلوحة تحكم لتوليد الكلام بالذكاء الاصطناعي ومولد برمجي لإنشاء مخرجات صوتية طبيعية. يركز المشروع على تخصيص ملفات تعريف الصوت والتحكم في فروق الكلام الدقيقة. يسمح بالحفاظ على خصائص متحدث متسقة باستخدام قيم البذور (Seeds) وملفات البيانات، مع توفير عناصر تحكم في النبرة والضحك والتوقفات من خلال مطالبات سلوكية ومعلمات أخذ العينات. يتضمن النظام معمارية عميل-خادم تتعامل مع معالجة الصوت غير المتزامنة وتوفر واجهة برمجية لتكامل التطبيقات الخارجية. يدير ملفات تعريف الصوت وتكوينات الصوت عبر واجهة مدارة الحالة لضمان توليد متسق.
Transforms mixed language text into spoken audio through a unified synthesis interface.
هذه مجموعة من النماذج العصبية المدربة مسبقاً للتعرف على الكلام، والتركيب، وكشف نشاط الصوت. توفر مكتبة من الأصول المصممة لتحويل الكلام إلى نص، وتحويل النص إلى كلام، وتحديد مقاطع الكلام البشري داخل الصوت. يتميز المشروع بتركيب النص إلى كلام مع دعم لغات متعددة واستخدام لغة ترميز تركيب الكلام للتحكم في النبرة، والطبقة، والتوقيت. بالنسبة للتعرف على الكلام، يتضمن النظام قدرات لنسخ الصوت إلى نص مع استخراج الطابع الزمني على مستوى الكلمة ومستعيد علامات ترقيم آلي لإدراج الأحرف الكبيرة وعلامات الترقيم في النص الخام. يتم تصدير النماذج إلى تنسيق Open Neural Network Exchange وTorchScript لتمكين التنفيذ عالي الأداء عبر مسرعات الأجهزة وأنظمة التشغيل المختلفة.
Provides synthesis models capable of generating speech across a wide variety of regional and minority languages.
Bark Voice Cloning is a text-to-speech synthesis engine designed to generate natural-sounding audio and replicate specific vocal characteristics. The system utilizes a transformer-based autoregressive model to convert written text into high-fidelity speech, supporting multilingual output and expressive delivery. The project distinguishes itself through zero-shot voice cloning, which extracts speaker identity embeddings from short audio samples to condition the generative model without requiring extensive fine-tuning. It also provides specialized workflows for voice identity conversion, allowi
Generates natural-sounding audio in multiple languages from text input using specialized phonetic models.
This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of conversational AI agents. It enables the creation of lifelike spoken audio from text, the replication of human voices through custom cloning, and the deployment of real-time voice agents capable of interacting with external large language models. The library distinguishes itself through deep integration of conversational AI capabilities, including the design of agent personas and the execution of real-time actions via APIs. It supports professional-grade audio production thro
Provides high-fidelity speech synthesis across multiple languages while maintaining native-level emotion and clarity.