5 مستودعات
Voice cloning systems that generate speech in multiple languages from reference samples.
Distinct from Voice Cloning Tools: Distinct from Voice Cloning Tools: specifically focuses on multilingual output and integration with multiple synthesis engines, not just cloning from a single sample.
Explore 5 awesome GitHub repositories matching artificial intelligence & ml · Multilingual Voice Cloning Synthesizers. Refine with filters or upvote what's useful.
wukong-robot is an open-source, Chinese-language voice assistant platform that integrates ChatGPT for multi-turn conversational AI. It is built around a plugin-based smart speaker framework, combining offline wake word detection with local speech synthesis to enable hands-free, voice-controlled interactions without requiring a constant internet connection. The platform distinguishes itself through its modular architecture, supporting custom wake word training via the command line and a plugin system that routes user intents using regular expressions for extensible functionality. It offers mul
Uses a voice model trained on a dataset to generate speech that mimics a specific person's voice.
Voice Pro is a comprehensive speech and audio processing toolkit that combines text-to-speech synthesis, voice cloning, speech recognition, and translation capabilities into a single application. At its core, the project enables users to generate natural-sounding speech from text, clone voices from short audio samples without requiring prior training data, and perform real-time speech translation across over 100 languages. The platform distinguishes itself through its integrated multimedia workflow, allowing users to download YouTube videos, extract audio, separate voice tracks, generate word
Generates speech using multiple cloning engines with support for celebrity voices and multilingual output.
This project is a TensorFlow voice conversion framework and deep learning audio toolkit designed for neural voice style transfer. It functions as a speech synthesis engine that transforms the spectral characteristics of a source speaker's voice to match the vocal identity of a target speaker. The system employs a phoneme-based approach to voice conversion, classifying audio utterances into speaker-independent phonemes and resynthesizing them using a target voice. This pipeline allows for the transformation of voice characteristics by mapping audio features between different speakers. The too
Transfers vocal characteristics using trained voice models to synthesize speech in a target identity.
Linly-Dubbing is an automated video dubbing pipeline designed for multilingual video localization. It converts spoken content in videos into another language by coordinating speech-to-text transcription, text translation, and text-to-speech synthesis. The system distinguishes itself through AI-driven lip synchronization and animation, which aligns facial expressions and mouth movements to the synthesized voiceover. It also utilizes audio source separation to isolate vocals from background music and noise, allowing for clean voice replacement while preserving original background audio. The br
Provides a voice synthesis system that supports voice cloning across multiple languages for video localization.
هذا التطبيق عبارة عن منصة لتوليف الصوت بالذكاء الاصطناعي واستنساخ الصوت العصبي. يوفر مجموعة أدوات شاملة لتحويل النص إلى كلام بشري يبدو طبيعياً من خلال تطبيق نماذج شبكة عصبية مدربة خصيصاً على عينات صوتية محددة. تسهل المنصة دورة حياة تطوير النموذج الصوتي بالكامل، بما في ذلك إعداد الكتب الصوتية الخام ونسخ الفيديو في مجموعات بيانات تدريب منظمة. كما تدعم تدريب هذه النماذج على أجهزة محلية أو بعيدة، باستخدام المعالجة الموزعة متعددة وحدات معالجة الرسومات (multi-GPU) للتعامل مع البيانات واسعة النطاق وتسريع تقارب النموذج. بعيداً عن التدريب، تتضمن المنصة قدرات لإدارة ونقل مجموعات البيانات الصوتية عبر بيئات تخزين مختلفة. يمكن للمستخدمين إجراء الاستدلال عن طريق ضبط المتغيرات الكامنة ومعلمات التوليف لتعديل النبرة، والتصريف العاطفي، والصفات الأسلوبية للمخرجات الصوتية المولدة. يعتمد التطبيق على تقنيات التعلم العميق لتحويل التمثيلات الصوتية إلى أشكال موجية عالية الدقة.
Converts text into natural human speech by applying trained neural network models to specific audio samples.