awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to as-ideas/transformertts

Open-source alternatives to TransformerTTS

30 open-source projects similar to as-ideas/transformertts, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best TransformerTTS alternative.

  • 2noise/chattts2noise avatar

    2noise/ChatTTS

    39,464View on GitHub↗

    ChatTTS is a conversational text-to-speech generative model designed to convert written dialogue into natural sounding audio. It functions as a multilingual speech synthesis framework capable of producing human-like audio across different languages and speaker profiles. The system is distinguished by its ability to generate interactive dialogue with realistic vocal nuances. It utilizes a speech nuance controller to insert specific tokens that trigger non-verbal elements, such as laughter, pauses, and interjections, during the synthesis process. The project includes a streaming audio generato

    Pythonagentchatchatgpt
    View on GitHub↗39,464
  • aishell-foundation/dacidianaishell-foundation avatar

    aishell-foundation/DaCiDian

    301View on GitHub↗

    DaCiDian is an open-sourced chinese mandarin lexicon for automatic speech recognition(ASR)

    Python
    View on GitHub↗301
  • aliutkus/speechmetricsaliutkus avatar

    aliutkus/speechmetrics

    1,051View on GitHub↗

    A wrapper around speech quality metrics MOSNet, BSSEval, STOI, PESQ, SRMR, SISDR

    Python
    View on GitHub↗1,051
  • boson-ai/higgs-audioboson-ai avatar

    boson-ai/higgs-audio

    7,919View on GitHub↗

    Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large language model architectures. It functions as a multilingual speech synthesizer capable of generating high-fidelity audio across different languages with control over emotional tone and prosody. The system includes a voice cloning tool that creates synthetic replicas of specific speakers from short audio samples without requiring extensive model training. It also provides a streaming audio API designed to deliver generated speech incrementally to minimize playback delay. The

    Python
    View on GitHub↗7,919

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • bytedance/megatts3bytedance avatar

    bytedance/MegaTTS3

    6,066View on GitHub↗

    MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English, including seamless code-switching within a single utterance. It functions as a text-to-speech engine, voice cloning system, and speech-to-text alignment tool, built around an acoustic latent compression model that encodes high-resolution audio into compact representations for efficient processing. The system distinguishes itself through accent intensity control, allowing adjustment of a speaker's accent strength in generated speech, and voice cloning from short audio samples for pers

    Pythonresearch
    View on GitHub↗6,066
  • emotional-text-to-speech/dl-for-emo-ttsEmotional-Text-to-Speech avatar

    Emotional-Text-to-Speech/dl-for-emo-tts

    458View on GitHub↗

    :computer: :robot: A summary on our attempts at using Deep Learning approaches for Emotional Text to Speech :speaker:

    Jupyter Notebook
    View on GitHub↗458
  • facebookresearch/covostfacebookresearch avatar

    facebookresearch/covost

    401View on GitHub↗

    CoVoST: A Large-Scale Multilingual Speech-To-Text Translation Corpus (CC0 Licensed)

    Python
    View on GitHub↗401
  • facebookresearch/omnilingual-asrfacebookresearch avatar

    facebookresearch/omnilingual-asr

    2,671View on GitHub↗

    Omnilingual-ASR is a multilingual automatic speech recognition framework and toolkit designed to transcribe audio across 1,600 languages. It provides a complete pipeline for converting speech to text, including a toolkit for fine-tuning pre-trained speech models to specific languages or datasets using custom training recipes. The system supports zero-shot speech recognition, allowing the model to predict text in unseen languages without extensive training data. It further enables few-shot language guidance through in-context examples and uses language codes to constrain transcription output t

    Python
    View on GitHub↗2,671
  • fighting41love/become-yukarinfighting41love avatar

    fighting41love/become-yukarin

    20View on GitHub↗

    Convert your voice to favorite voice

    View on GitHub↗20
  • fireredteam/fireredtts2F

    FireRedTeam/FireRedTTS2

    0View on GitHub↗
    View on GitHub↗0
  • fishaudio/fish-speechfishaudio avatar

    fishaudio/fish-speech

    24,928View on GitHub↗

    This project is a generative speech synthesis engine that converts text into high-fidelity human speech. It utilizes a two-stage autoregressive transformer architecture that separates semantic token prediction from acoustic detail reconstruction to balance linguistic accuracy with audio quality. The system is designed to support multilingual output and conversational AI development, enabling the generation of context-aware speech that maintains flow across multiple dialogue turns. The platform distinguishes itself through a production-ready inference server that employs continuous batching to

    Pythonllamatransformertts
    View on GitHub↗24,928
  • funaudiollm/cosyvoiceFunAudioLLM avatar

    FunAudioLLM/CosyVoice

    21,673View on GitHub↗

    CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus

    Pythonaudio-generationcantonesechatbot
    View on GitHub↗21,673
  • google/visqolgoogle avatar

    google/visqol

    896View on GitHub↗

    Perceptual Quality Estimator for speech and audio

    C++
    View on GitHub↗896
  • hexgrad/kokorohexgrad avatar

    hexgrad/kokoro

    5,729View on GitHub↗

    Kokoro is a lightweight neural text-to-speech engine that converts written text into spoken audio using a compact model designed for fast inference. It supports multiple languages through language-specific grapheme-to-phoneme conversion pipelines, and offers voice profile selection to change the character of the generated speech. The engine provides GPU acceleration on Apple Silicon hardware by setting a single environment variable, enabling faster inference on Mac M-series machines. It also includes pattern-based text segmentation, allowing input text to be split at user-defined delimiters t

    JavaScript
    View on GitHub↗5,729
  • ideo/laughdetectionideo avatar

    ideo/LaughDetection

    130View on GitHub↗
    Jupyter Notebook
    View on GitHub↗130
  • inclusionai/ming-omni-ttsI

    inclusionAI/Ming-omni-tts

    0View on GitHub↗
    View on GitHub↗0
  • index-tts/index-ttsindex-tts avatar

    index-tts/index-tts

    18,851View on GitHub↗

    Index-tts is a neural audio generation engine designed to convert written text into high-fidelity human speech. By utilizing deep learning models and phoneme-based sequence modeling, the system transforms text into natural-sounding audio waveforms suitable for a variety of accessibility and media applications. The platform functions as a server-side inference pipeline that provides a programmatic interface for integrating voice generation into external applications. It distinguishes itself through asynchronous audio streaming, which buffers and delivers generated speech chunks in real time to

    Pythonbigvgancross-lingualindextts
    View on GitHub↗18,851
  • iver56/audiomentationsiver56 avatar

    iver56/audiomentations

    2,286View on GitHub↗

    A Python library for audio data augmentation. Useful for making audio ML models work well in the real world, not just in the lab.

    Python
    View on GitHub↗2,286
  • k2-fsa/omnivoicek2-fsa avatar

    k2-fsa/OmniVoice

    7,689View on GitHub↗

    OmniVoice is a Python-based open-source project hosted under the k2-fsa organization. The repository currently does not contain any documented features or structured documentation, making it difficult to determine its specific purpose or capabilities at this time.

    Python
    View on GitHub↗7,689
  • k2-fsa/zipvoiceK

    k2-fsa/ZipVoice

    0View on GitHub↗
    View on GitHub↗0
  • kittenml/kittenttsKittenML avatar

    KittenML/KittenTTS

    10,044View on GitHub↗

    KittenTTS is a neural text-to-speech engine and text-to-audio synthesis tool that converts written text into spoken audio using lightweight neural network models. It functions as both a speech synthesizer and an audio file generator, producing spoken audio for offline playback. The system includes a text normalization processor that expands numbers and abbreviations into full spoken words to improve the naturalness of the synthesized speech. It supports diverse voice options and provides the ability to adjust playback speed.

    Python
    View on GitHub↗10,044
  • kuangdd/aukitK

    KuangDD/aukit

    0View on GitHub↗
    View on GitHub↗0
  • kuangdd/phkitK

    KuangDD/phkit

    0View on GitHub↗
    View on GitHub↗0
  • kuangdd/zhrtvcK

    KuangDD/zhrtvc

    0View on GitHub↗
    View on GitHub↗0
  • kuangdd/zhvoiceK

    KuangDD/zhvoice

    0View on GitHub↗
    View on GitHub↗0
  • kyutai-labs/delayed-streams-modelingkyutai-labs avatar

    kyutai-labs/delayed-streams-modeling

    2,955View on GitHub↗

    Kyutai's Speech-To-Text and Text-To-Speech models based on the Delayed Streams Modeling framework.

    Python
    View on GitHub↗2,955
  • lukhy/masrlukhy avatar

    lukhy/masr

    1,969View on GitHub↗

    中文语音识别; Mandarin Automatic Speech Recognition;

    Python
    View on GitHub↗1,969
  • microsoft/vibevoicemicrosoft avatar

    microsoft/VibeVoice

    49,394View on GitHub↗

    VibeVoice is a generative artificial intelligence platform designed for text-to-speech synthesis. It functions as a neural audio generation framework that converts written text into natural-sounding spoken audio, specifically engineered to maintain consistent vocal characteristics and narrative prosody across extended passages of content. The system distinguishes itself through its ability to generate long-form conversational speech while preserving speaker identity and linguistic content. By utilizing latent space disentanglement, the model separates speaker traits from the input text, allow

    Python
    View on GitHub↗49,394
  • midas-research/audinomidas-research avatar

    midas-research/audino

    1,140View on GitHub↗

    Open source audio annotation tool for humans

    TypeScript
    View on GitHub↗1,140
  • miteshputhranneu/speech-emotion-analyzerMITESHPUTHRANNEU avatar

    MITESHPUTHRANNEU/Speech-Emotion-Analyzer

    1,408View on GitHub↗

    The neural network model is capable of detecting five different male/female emotions from audio speeches. (Deep Learning, NLP, Python)

    Jupyter Notebook
    View on GitHub↗1,408