awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

6 रिपॉजिटरी

Awesome GitHub RepositoriesMultilingual Text-to-Speech Engines

Text-to-speech engines that support synthesis across multiple languages with automatic language detection and fallback.

Distinct from Text-to-Speech Conversions: Distinct from Text-to-Speech Conversions: specifically handles multilingual output with language detection, not single-language conversion.

Explore 6 awesome GitHub repositories matching artificial intelligence & ml · Multilingual Text-to-Speech Engines. Refine with filters or upvote what's useful.

Awesome Multilingual Text-to-Speech Engines GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • myshell-ai/melottsmyshell-ai का अवतार

    myshell-ai/MeloTTS

    7,509GitHub पर देखें↗

    MeloTTS is an open-source text-to-speech library that generates natural-sounding speech across six languages, with the ability to mix two languages within a single utterance. Its architecture combines a token-based text frontend with a language-agnostic acoustic model, enabling it to handle bilingual code-switching and produce streaming audio output in real time. The system is designed to run efficiently on standard CPU hardware without requiring a dedicated GPU, using a lightweight neural network for real-time inference. It supports English, Spanish, French, Chinese, Japanese, and Korean, an

    Converts written text into natural-sounding speech across multiple languages including English, Spanish, French, Chinese, Japanese, and Korean.

    Pythonchineseenglishfrench
    GitHub पर देखें↗7,509
  • zyphra/zonosZyphra का अवतार

    Zyphra/Zonos

    7,225GitHub पर देखें↗

    Zonos is a controllable audio synthesis engine and large language model for text-to-speech. It serves as a multilingual speech generator capable of producing audio in English, Japanese, Chinese, French, and German. The system provides zero-shot voice cloning, allowing the replication of specific human voices using short audio samples. It supports the capture of nuanced behaviors, such as whispering, and provides parametric control over speaking rate, pitch, frequency, and emotional tone. The project covers a broad range of expressive speech synthesis and custom audio generation capabilities,

    Provides a multilingual text-to-speech engine supporting multiple languages for global accessibility.

    Python
    GitHub पर देखें↗7,225
  • espeak-ng/espeak-ngespeak-ng का अवतार

    espeak-ng/espeak-ng

    6,604GitHub पर देखें↗

    espeak-ng एक बहुभाषी टेक्स्ट-टू-स्पीच इंजन और C-आधारित लाइब्रेरी है जो लिखित टेक्स्ट को विभिन्न भाषाओं, लहजों और क्षेत्रीय बोलियों में बोले गए ऑडियो में परिवर्तित करती है। यह बाहरी एप्लिकेशन में संश्लेषण क्षमताओं को एम्बेड करने के लिए एक प्रोग्रामेटिक इंटरफेस और एक फोनेटिक टेक्स्ट कनवर्टर दोनों के रूप में कार्य करता है जो लिखित टेक्स्ट को फोनेम कोड में अनुवादित करता है। यह सिस्टम कई संश्लेषण विधियों का उपयोग करता है, जिसमें गणितीय रूप से मुखर ध्वनियाँ उत्पन्न करने के लिए फॉर्मेंट संश्लेषण और पूर्व-रिकॉर्ड किए गए फोनेटिक सेगमेंट को जोड़कर ऑडियो बनाने के लिए डिफोन संश्लेषण शामिल है। इसमें ऑडियो पिच और टाइमिंग को नियंत्रित करने के लिए SSML और HTML टैग को पार्स करने में सक्षम एक स्पीच प्रोसेसर शामिल है। इंजन फोनेटिक अनुवाद मैप्स और परिभाषा फ़ाइलों के माध्यम से कस्टम वॉयस डिज़ाइन और भाषा उच्चारण अनुकूलन के लिए टूल्स प्रदान करता है। यह WAV प्रारूप में ऑडियो फ़ाइल एक्सपोर्ट, प्लेबैक गति समायोजन और भाषाई विश्लेषण के लिए फोनेटिक डेटा के निर्माण का समर्थन करता है। स्पीच जनरेशन को ट्रिगर करने और ऑडियो आउटपुट सेटिंग्स को प्रबंधित करने के लिए एक कमांड लाइन इंटरफेस उपलब्ध है।

    Acts as a comprehensive engine for converting written text into spoken audio across various languages and dialects.

    C
    GitHub पर देखें↗6,604
  • plachtaa/vits-fast-fine-tuningPlachtaa का अवतार

    Plachtaa/VITS-fast-fine-tuning

    5,016GitHub पर देखें↗

    VITS-fast-fine-tuning छोटे ऑडियो डेटासेट का उपयोग करके विशिष्ट टारगेट आवाज़ों के लिए स्पीच सिंथेसिस मॉडल्स को अनुकूलित करने के लिए एक पाइपलाइन है। यह एक तेज़ स्पीकर अनुकूलन टूल और एक बहुभाषी स्पीच सिंथेसाइज़र के रूप में कार्य करता है जो विभिन्न भाषाओं में बोले गए ऑडियो को जनरेट करने में सक्षम है। यह सिस्टम मेनी-टू-मेनी वॉयस कन्वर्ज़न के लिए एक फ़्रेमवर्क प्रदान करता है, जो मूल भाषाई सामग्री को संरक्षित करते हुए एक स्पीकर की पहचान को दूसरे में बदल देता है। यह ऑडियो क्लिप्स या वीडियो स्रोतों के साथ एक प्री-ट्रेंड मॉडल को फ़ाइन-ट्यून करके टेक्स्ट-टू-स्पीच के लिए आवाज़ के अनुकूलन की अनुमति देता है। यह प्रोजेक्ट एंड-टू-एंड स्पीच सिंथेसिस और ऑडियो प्रोसेसिंग को कवर करता है, जो उच्च-निष्ठा (high-fidelity) ऑडियो उत्पन्न करने के लिए एडवरसैरियल वेवफ़ॉर्म जनरेशन और मोनोटोनिक अलाइनमेंट सर्च का उपयोग करता है। यह बोलने की लय में विविधताओं को प्रबंधित करने के लिए एक स्टोकेस्टिक ड्यूरेशन प्रेडिक्टर को शामिल करता है और प्री-ट्रेंड मॉडल ट्रांसफर का समर्थन करता है।

    Supports synthesis of spoken audio across multiple languages while maintaining consistent character voices.

    Python
    GitHub पर देखें↗5,016
  • whisperspeech/whisperspeechWhisperSpeech का अवतार

    WhisperSpeech/WhisperSpeech

    4,617GitHub पर देखें↗

    WhisperSpeech एक बहुभाषी स्पीच सिंथेसाइज़र और न्यूरल टेक्स्ट-टू-स्पीच सिस्टम है। यह टेक्स्ट को हाई-फिडेलिटी सिंथेटिक ऑडियो में बदलने के लिए Whisper मॉडल आर्किटेक्चर को इनवर्ट करके कार्य करता है। यह सिस्टम विशिष्ट वक्ताओं की नकल करने के लिए संदर्भ ऑडियो फ़ाइलों का उपयोग करके वॉयस क्लोनिंग को सक्षम बनाता है। यह बहुभाषी स्पीच प्रोडक्शन का समर्थन करता है, जिसमें विभिन्न भाषाओं में ऑडियो उत्पन्न करने और एक ही वाक्य के भीतर भाषा स्विचिंग को संभालने की क्षमता शामिल है। यह प्रोजेक्ट टेक्स्ट-टू-स्पीच जनरेशन और स्पीच डेटासेट तैयारी सहित स्पीच क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। इसमें स्पीच को टेक्स्ट में ट्रांसक्राइब करने, ध्वनिक टोकन निकालने, और वॉयस एक्टिविटी का पता लगाने के लिए टूल्स शामिल हैं।

    Specializes in multilingual speech production with seamless mixing of multiple languages in one output.

    Jupyter Notebookpytorchspeech-synthesistts
    GitHub पर देखें↗4,617
  • supertone-inc/supertonicsupertone-inc का अवतार

    supertone-inc/supertonic

    2,626GitHub पर देखें↗

    Supertonic is an on-device neural text-to-speech engine that runs entirely locally without cloud dependencies or GPU acceleration. It converts written text into natural-sounding speech across 31 languages with automatic language detection and a fallback model for unsupported locales. The engine provides expressive speech control through inline prosody tags that dynamically adjust pitch, rate, and tone during synthesis. It supports voice cloning from a short reference audio clip by extracting a speaker embedding vector, and offers a selection of pre-built voices tuned for different use cases.

    Converts written text into natural-sounding speech across 31 languages with automatic language detection.

    C++cppcsharpgo
    GitHub पर देखें↗2,626
  1. Home
  2. Artificial Intelligence & ML
  3. Speech and Text Conversion
  4. Text-to-Speech Conversions
  5. Multilingual Text-to-Speech Engines