awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com

Open-Source Alternatives to ElevenLabs

Ranking updated Aug 19, 2026

For an open source voice cloning and speech generation tool, the strongest matches are boson-ai/higgs-audio (Higgs-audio is a neural text-to-speech engine built in Python), babysor/mockingbird (MockingBird is a self-hostable neural text-to-speech engine built in) and mozilla/tts (This project provides a comprehensive neural text-to-speech engine and). corentinj/real-time-voice-cloning and myshell-ai/openvoice round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

We curate open-source GitHub repositories matching “open source alternatives to elevenlabs”. Results are ranked by relevance to your query — pick filters below to narrow, or refine with AI.

Open-Source Alternatives to ElevenLabs

Find the best repos with AI.We'll search the best matching repositories with AI.
  • boson-ai/higgs-audioboson-ai avatar

    boson-ai/higgs-audio

    7,919View on GitHub↗

    Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large language model architectures. It functions as a multilingual speech synthesizer capable of generating high-fidelity audio across different languages with control over emotional tone and prosody. The system includes a voice cloning tool that creates synthetic replicas of specific speakers from short audio samples without requiring extensive model training. It also provides a streaming audio API designed to deliver generated speech incrementally to minimize playback delay. The

    Higgs-audio is a neural text-to-speech engine built in Python that supports zero-shot voice cloning, multilingual generation, streaming audio APIs, and pretrained large language model architectures for self-hosted conversational voice synthesis.

    PythonMultilingual Speech ModelsNeural Text-to-Speech EnginesVoice Cloning
    View on GitHub↗7,919
  • babysor/mockingbirdbabysor avatar

    babysor/MockingBird

    36,903View on GitHub↗

    MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions as a voice synthesis trainer for building custom models from audio datasets, a command-line generator for producing audio files, and a text-to-speech server for remote application integration. The project specializes in real-time voice cloning, which extracts vocal characteristics from short audio samples to mimic a target speaker's unique timbre. It utilizes reference-driven audio synthesis to condition pre-trained models on specific audio samples, allowing for the generation

    MockingBird is a self-hostable neural text-to-speech engine built in Python that features real-time voice cloning from short audio samples, custom model training, and a server for integration.

    PythonNeural Text-to-Speech EnginesVoice CloningVoice Cloning Tools
    View on GitHub↗36,903
  • mozilla/ttsmozilla avatar

    mozilla/TTS

    10,151View on GitHub↗

    This project is a comprehensive suite for neural speech synthesis, featuring a deep learning text-to-speech engine, a neural speech synthesis trainer, and a voice cloning toolkit. It provides a system for synthesizing human-like speech from text using neural network models and high-fidelity vocoders. The suite includes a speech model conversion utility to transform deep learning models between different formats for deployment across various hardware runtimes. It also provides a self-contained HTTP server to expose pre-trained text-to-speech models as a remote audio API. Capabilities include

    This project provides a comprehensive neural text-to-speech engine and voice cloning toolkit complete with pre-trained models, a Python API, and a self-hosted server for speech synthesis.

    Jupyter NotebookNeural Text-to-Speech EnginesVoice CloningVoice Cloning Toolkits
    View on GitHub↗10,151
  • corentinj/real-time-voice-cloningCorentinJ avatar

    CorentinJ/Real-Time-Voice-Cloning

    59,918View on GitHub↗

    This project is a neural text-to-speech engine and voice cloning toolkit designed to generate synthetic speech that mimics the vocal characteristics of a target speaker. It functions as a real-time audio synthesizer, utilizing a deep learning pipeline to convert written text into high-fidelity speech output with minimal latency. The system employs a transfer learning framework that leverages pre-trained speaker verification models to adapt synthesis to new, unseen vocal identities. By using an encoder-based speaker embedding process, the toolkit maps variable-length audio samples into a laten

    This repository is a neural text-to-speech engine and voice cloning toolkit that implements speaker embeddings and transfer learning for synthetic speech generation in Python, fitting your search for a self-hostable voice generation system.

    PythonNeural Text-to-Speech EnginesVoice Cloning ToolsReal-Time Voice Cloning
    View on GitHub↗59,918
  • myshell-ai/openvoicemyshell-ai avatar

    myshell-ai/OpenVoice

    36,720View on GitHub↗

    OpenVoice is a multilingual text-to-speech framework and voice cloning AI model designed for high-fidelity voice replication and low-latency audio generation. It functions as an instant speech synthesis engine that converts text to audio while replicating a specific speaker's tone and color. The system is distinguished by its ability to perform cross-lingual cloning, allowing the vocal characteristics of a reference speaker to be applied to speech in different languages regardless of the original training data. It utilizes a decoupled representation to separate the physical identity of a voic

    OpenVoice is a self-hostable neural text-to-speech framework that provides zero-shot voice cloning, cross-lingual voice transfer, and a Python API for high-fidelity speech generation.

    PythonNeural Text-to-Speech EnginesVoice CloningZero-Shot Voice Cloning
    View on GitHub↗36,720
  • zyphra/zonosZyphra avatar

    Zyphra/Zonos

    7,225View on GitHub↗

    Zonos is a controllable audio synthesis engine and large language model for text-to-speech. It serves as a multilingual speech generator capable of producing audio in English, Japanese, Chinese, French, and German. The system provides zero-shot voice cloning, allowing the replication of specific human voices using short audio samples. It supports the capture of nuanced behaviors, such as whispering, and provides parametric control over speaking rate, pitch, frequency, and emotional tone. The project covers a broad range of expressive speech synthesis and custom audio generation capabilities,

    Zonos is an open-source controllable audio synthesis and text-to-speech engine featuring zero-shot voice cloning, multilingual support, and rich prosody controls that directly match your self-hosted generation requirements.

    PythonVoice CloningVoice Cloning EnginesZero-Shot Voice Cloning
    View on GitHub↗7,225
  • coqui-ai/ttscoqui-ai avatar

    coqui-ai/TTS

    45,568View on GitHub↗

    This project is a deep learning text-to-speech toolkit used for training and deploying neural speech synthesis models. It provides a comprehensive framework for converting written text into spoken audio, utilizing neural vocoders to transform synthesized spectrograms into high-fidelity audio waveforms. The toolkit includes a voice cloning system that replicates specific human voices by extracting speaker embeddings from short audio samples. It also supports multi-speaker audio synthesis, allowing the generation of speech across different vocal identities using specialized model architectures.

    This open-source deep learning toolkit provides neural text-to-speech synthesis, voice cloning from audio samples, pretrained models, and a Python API for developers.

    PythonNeural Text-to-Speech EnginesVoice CloningMulti-Speaker Synthesis
    View on GitHub↗45,568
  • swivid/f5-ttsSWivid avatar

    SWivid/F5-TTS

    14,798View on GitHub↗

    F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent synthetic speech. It functions as a multilingual speech synthesizer and neural training framework, providing tools for voice cloning and high-performance inference serving. The project distinguishes itself through a voice cloning toolkit capable of mimicking specific speaker characteristics and tones from reference audio clips. It supports cross-lingual generation, allowing for the synthesis of audio across various global languages or the mixing of multiple languages within a s

    F5-TTS is a self-hostable neural text-to-speech framework providing pretrained models, streaming audio, multi-lingual support, and a Python API for voice cloning from audio samples.

    PythonCross-Lingual Speech GeneratorsVoice CloningVoice Cloning Engines
    View on GitHub↗14,798
  • funaudiollm/cosyvoiceFunAudioLLM avatar

    FunAudioLLM/CosyVoice

    21,673View on GitHub↗

    CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus

    CosyVoice is a neural text-to-speech engine and speech synthesis framework that delivers zero-shot voice cloning, multilingual support, and expressive audio generation via a Python-friendly architecture.

    PythonMultilingual Speech ModelsNeural Text-to-Speech EnginesZero-Shot Voice Cloning
    View on GitHub↗21,673
  • koljab/realtimettsKoljaB avatar

    KoljaB/RealtimeTTS

    3,964View on GitHub↗

    RealtimeTTS is a real-time text-to-speech engine and stream processor designed to convert text or token streams into audio playback with minimal latency. It provides a programmatic interface for managing audio streams, synthesis progress, and the integration of local or cloud-based speech engines. The system includes a neural voice cloning tool that generates synthetic speech by extracting acoustic features from reference audio samples. It utilizes a provider-based abstraction to route synthesis requests across different neural models and cloud APIs. The project covers a range of functional

    RealtimeTTS is a real-time text-to-speech engine and stream processor that supports voice cloning from audio samples and integrates with local or cloud-based neural models through a Python API, though it focuses more on streaming and playback orchestration than serving as an all-in-one pretrained standalone speech generator.

    PythonReal-Time Speech SynthesisVoice Cloning EnginesVoice Cloning Tools
    View on GitHub↗3,964
  • rvc-boss/gpt-sovitsRVC-Boss avatar

    RVC-Boss/GPT-SoVITS

    58,724View on GitHub↗

    GPT-SoVITS is a text-to-speech synthesis engine and voice cloning toolkit designed for generating natural-sounding human speech. It functions as a neural audio processing pipeline that maps input text to high-fidelity audio waveforms, utilizing conditional variational autoencoders and flow-based decoders to ensure expressive output. The platform distinguishes itself through its ability to perform few-shot voice cloning and cross-lingual speech generation, allowing users to maintain a specific speaker's vocal identity and emotional delivery across multiple languages. By employing cross-modal l

    GPT-SoVITS is a self-hostable neural text-to-speech engine and voice cloning toolkit with a Python API, supporting cross-lingual generation and few-shot voice adaptation from audio samples.

    PythonCross-Lingual Speech GeneratorsVoice Cloning Tools
    View on GitHub↗58,724
  • sparkaudio/spark-ttsSparkAudio avatar

    SparkAudio/Spark-TTS

    10,930View on GitHub↗

    Spark-TTS is a deep learning text-to-speech synthesis engine designed to convert written text into high-fidelity audio. It utilizes a transformer-based architecture and autoregressive sequence modeling to generate coherent speech, transforming linguistic input into natural-sounding waveforms through neural speech codec synthesis. The platform distinguishes itself through zero-shot voice cloning, which allows users to mimic a target speaker’s unique vocal identity using only a short reference audio sample without requiring additional model training. It also features cross-lingual phonetic mapp

    Spark-TTS is an autoregressive deep learning text-to-speech engine that supports zero-shot voice cloning and cross-lingual capabilities, fulfilling most of your requirements though lacking explicit mention of streaming audio or pre-packaged developer APIs.

    PythonCross-Lingual Speech GeneratorsMultilingual Speech ModelsZero-Shot Voice Cloning
    View on GitHub↗10,930
  • fishaudio/fish-speechfishaudio avatar

    fishaudio/fish-speech

    24,928View on GitHub↗

    This project is a generative speech synthesis engine that converts text into high-fidelity human speech. It utilizes a two-stage autoregressive transformer architecture that separates semantic token prediction from acoustic detail reconstruction to balance linguistic accuracy with audio quality. The system is designed to support multilingual output and conversational AI development, enabling the generation of context-aware speech that maintains flow across multiple dialogue turns. The platform distinguishes itself through a production-ready inference server that employs continuous batching to

    Fish Speech is an open-source generative speech synthesis engine providing voice cloning, neural text-to-speech, multilingual support, and a Python API for self-hosted developer integration.

    PythonMultilingual Speech ModelsVoice CloningVoice Cloning Toolkits
    View on GitHub↗24,928
  • plachtaa/vall-e-xPlachtaa avatar

    Plachtaa/VALL-E-X

    7,939View on GitHub↗

    VALL-E-X is a neural speech synthesis framework and zero-shot text-to-speech engine. It functions as a multilingual synthesizer capable of generating natural human speech with control over emotion, pitch, and prosody. The project specializes in zero-shot voice cloning and cross-lingual voice replication, allowing the system to produce personalized speech in multiple target languages using short audio samples without additional training. It further enables cross-language accent manipulation and the ability to match the emotional tone and acoustic environment of a provided prompt. The implemen

    This repository is a neural speech synthesis framework that provides zero-shot voice cloning and multilingual text-to-speech generation in Python, though it lacks some deployment features like streaming audio and built-in self-hosting wrappers.

    PythonCross-Lingual Speech GeneratorsMultilingual Speech ModelsVoice Cloning
    View on GitHub↗7,939
  • neonbjb/tortoise-ttsneonbjb avatar

    neonbjb/tortoise-tts

    14,864View on GitHub↗

    Tortoise-tts is a neural text-to-speech engine and voice cloning toolkit designed for high-quality audio generation. It functions as a zero-shot synthesis system, meaning it can generate speech for unseen speakers without requiring additional training or fine-tuning for each new voice. The system specializes in replicating human vocal characteristics using small sets of reference audio clips. It allows for the extraction of voice latents to mimic specific speakers, the generation of random synthetic identities, and the blending of multiple voice profiles to create hybrid vocal identities. Th

    Tortoise-tts is a neural text-to-speech engine and zero-shot voice cloning toolkit that specializes in generating realistic human speech from reference audio clips.

    Jupyter NotebookVoice CloningZero-Shot Voice CloningVoice Cloning Toolkits
    View on GitHub↗14,864
  • openbmb/voxcpmOpenBMB avatar

    OpenBMB/VoxCPM

    29,985View on GitHub↗

    VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice cloning tool and a synthetic voice designer, capable of generating natural speech across global languages and regional dialects using a GPU-accelerated audio generator. The project features a speech model fine-tuning framework that supports both full parameter updates and low-rank adaptation for customizing voice characteristics. It enables high-fidelity voice cloning from reference audio, including cross-lingual voice transfer and acoustic environment mimicry, as well as the crea

    VoxCPM is a multilingual speech synthesis system and text-to-speech inference server that supports voice cloning from reference audio, self-hosted deployment, and fine-tuning via a Python-friendly toolkit.

    PythonMultilingual Speech ModelsVoice CloningZero-Shot Voice Cloning
    View on GitHub↗29,985
  • metavoiceio/metavoice-srcmetavoiceio avatar

    metavoiceio/metavoice-src

    4,202View on GitHub↗

    This project is an expressive text-to-speech foundation model and voice cloning system designed to synthesize human-like speech with emotional nuance and high fidelity. It functions as a finetunable speech model that can generate audio mimicking a specific person using a reference voice sample. The system distinguishes itself through a high-performance inference engine that utilizes memory caching and hardware compilation to reduce latency during the audio generation process. It further allows for synthesis quality improvements by training the language model on custom datasets consisting of a

    This project is an expressive text-to-speech foundation model and voice cloning system built in Python, making it a strong fit for self-hosted neural speech generation, though it lacks some specific features like streaming audio and multi-lingual support mentioned in the search.

    PythonVoice CloningVoice Cloning EnginesZero-Shot Voice Cloning
    View on GitHub↗4,202
  • yl4579/styletts2yl4579 avatar

    yl4579/StyleTTS2

    6,294View on GitHub↗

    StyleTTS2 is an adversarial text-to-speech model that uses style diffusion and large speech language models to generate natural-sounding speech from text input. It combines adversarial training with large pre-trained speech models to improve speech quality and reduce artifacts, while employing a style diffusion process that extracts prosodic and timbral features from reference audio to guide speech generation. The model supports multi-speaker voice synthesis by conditioning the diffusion process on speaker-specific embeddings derived from reference utterances, enabling voice cloning and adapt

    StyleTTS2 is a self-hostable neural text-to-speech engine built in Python that supports few-shot voice cloning from reference audio samples, advanced adversarial training, and pretrained model loading.

    PythonFew-Shot Voice CloningMulti-Speaker SynthesisSpeaker Embeddings
    View on GitHub↗6,294
  • kevinwang676/bark-voice-cloningKevinWang676 avatar

    KevinWang676/Bark-Voice-Cloning

    2,957View on GitHub↗

    Bark Voice Cloning is a text-to-speech synthesis engine designed to generate natural-sounding audio and replicate specific vocal characteristics. The system utilizes a transformer-based autoregressive model to convert written text into high-fidelity speech, supporting multilingual output and expressive delivery. The project distinguishes itself through zero-shot voice cloning, which extracts speaker identity embeddings from short audio samples to condition the generative model without requiring extensive fine-tuning. It also provides specialized workflows for voice identity conversion, allowi

    This project provides a text-to-speech synthesis engine featuring zero-shot voice cloning and multilingual output based on autoregressive models, though it is packaged as a notebook implementation rather than a ready-to-deploy standalone server.

    Jupyter NotebookVoice CloningVoice Cloning ToolsZero-Shot Voice Cloning
    View on GitHub↗2,957
  • whisperspeech/whisperspeechWhisperSpeech avatar

    WhisperSpeech/WhisperSpeech

    4,617View on GitHub↗

    WhisperSpeech is a multilingual speech synthesizer and neural text-to-speech system. It functions by inverting the Whisper model architecture to convert text into high-fidelity synthetic audio. The system enables voice cloning by using reference audio files to mimic specific speakers. It supports multilingual speech production, which includes the ability to generate audio across different languages and handle language switching within a single sentence. The project covers a broad range of speech capabilities, including text-to-speech generation and speech dataset preparation. It incorporates

    WhisperSpeech is a neural text-to-speech engine and multilingual speech synthesizer written in Python that supports voice cloning from reference audio, though it lacks built-in streaming audio support.

    Jupyter NotebookVoice Cloning EnginesVoice Cloning ToolsZero-Shot Voice Cloning
    View on GitHub↗4,617
  • jasonppy/voicecraftjasonppy avatar

    jasonppy/VoiceCraft

    8,500View on GitHub↗

    VoiceCraft is a neural speech generation and manipulation system consisting of a text-to-speech system, a voice cloning tool, and an audio inpainting engine. It uses a large language model approach to synthesize high-fidelity audio from text and replicate speaker identities. The system provides zero-shot voice cloning and speech editing capabilities, allowing users to modify spoken content within existing recordings. This includes an audio inpainting engine that replaces specific sections of audio with new speech while preserving the original acoustic characteristics and speaker identity. Th

    VoiceCraft is an open-source neural speech generation and voice cloning system capable of zero-shot synthesis and speech editing, making it a strong fit for AI text-to-speech tasks though it lacks explicit streaming or packaged deployment features.

    Jupyter NotebookVoice CloningVoice Cloning ToolsZero-Shot Voice Cloning
    View on GitHub↗8,500
  • jianchang512/clone-voicejianchang512 avatar

    jianchang512/clone-voice

    8,959View on GitHub↗

    This project is a GPU-accelerated speech engine and AI voice cloning tool. It functions as a text-to-speech synthesizer and voice-to-voice converter that replicates specific human voices to generate synthetic speech. The system creates digital voice profiles by analyzing short audio samples or capturing live microphone input. These profiles enable the transformation of existing audio recordings into a target speaker's voice or the synthesis of new audio from written text. The engine supports subtitle-based speech generation for batch processing and automated dubbing workflows. A web-based au

    This repository provides a GPU-accelerated neural text-to-speech engine with voice cloning and profile management, though it lacks explicit mention of streaming audio support among its core features.

    PythonNeural Text-to-Speech EnginesVoice CloningVoice Cloning Tools
    View on GitHub↗8,959
  • canopyai/orpheus-ttscanopyai avatar

    canopyai/Orpheus-TTS

    6,201View on GitHub↗

    Orpheus-TTS is an open-source text-to-speech system that generates human-like audio with controllable emotional tone and the ability to clone voices from short audio samples. It is built on an architecture that treats speech generation as a language modeling task, using a large language model trained on text-speech pairs to produce audio tokens autoregressively. The system distinguishes itself through several key capabilities. It supports emotion-controllable speech synthesis by embedding emotional and intonation markers directly into text prompts, allowing the model to condition its output o

    Orpheus-TTS is a self-hostable text-to-speech engine built for voice cloning and emotional synthesis from audio samples using an autoregressive language model, matching the core requirements.

    PythonLow-Latency Audio StreamsVoice Cloning EnginesZero-Shot Voice Cloning
    View on GitHub↗6,201
  • nari-labs/dianari-labs avatar

    nari-labs/dia

    19,324View on GitHub↗

    Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt

    Dia is a text-to-speech and voice cloning engine built for production deployment and custom voice generation, though it lacks explicit mentions of streaming audio or multi-lingual support in the provided evidence.

    PythonNeural Text-to-Speech EnginesVoice CloningVoice Cloning Engines
    View on GitHub↗19,324
  • kyutai-labs/pocket-ttskyutai-labs avatar

    kyutai-labs/pocket-tts

    3,301View on GitHub↗

    Pocket-tts is a text-to-speech server and neural speech synthesizer that converts written text into audible speech. It includes a CPU-optimized inference engine and a voice cloning tool capable of analyzing audio samples to reproduce specific speaker characteristics. The system differentiates itself through the use of dynamic int8 quantization to reduce memory usage and increase generation speed on processors. It supports real-time speech synthesis by streaming audio chunks incrementally and utilizes voice state caching to store processed embeddings as portable files, bypassing redundant proc

    Pocket-tts provides a CPU-optimized text-to-speech server with voice cloning and real-time streaming, though it lacks explicit mention of multi-lingual support and pretrained model variety out of the box.

    PythonReal-Time Speech SynthesisVoice CloningVoice Cloning Tools
    View on GitHub↗3,301
  • neuphonic/neuttsneuphonic avatar

    neuphonic/neutts

    6,007View on GitHub↗

    Neutts is a neural text-to-speech engine designed for real-time streaming output on edge devices such as phones and laptops. It supports voice cloning from short audio references, enabling zero-shot reproduction of a target speaker's voice, and can be fine-tuned or retrained from scratch for custom voices and styles. The system distinguishes itself through a decoder-only architecture that halves memory and accelerates generation on constrained hardware, combined with quantized model inference for reduced memory footprint. Its streaming decoder loop interleaves synthesis with playback, deliver

    Neutts is a neural text-to-speech engine featuring real-time streaming and voice cloning from short audio references, though it lacks explicit multi-lingual support mentioned in the metadata.

    PythonNeural Text-to-Speech EnginesVoice CloningVoice Cloning Engines
    View on GitHub↗6,007
  • openbmb/minicpm-oOpenBMB avatar

    OpenBMB/MiniCPM-o

    23,850View on GitHub↗

    MiniCPM-o is a multimodal large language model designed to function as a real-time conversational assistant on edge devices. By mapping text, image, video, and audio inputs into a unified latent space, the system enables simultaneous cross-modal reasoning and full-duplex interaction. It is built as an edge-side inference engine, utilizing quantized model weights to maintain high-performance processing on consumer hardware. The system distinguishes itself through its integrated speech synthesis and voice cloning capabilities, which allow for the generation of expressive, personalized vocal out

    This repository provides a multimodal model with zero-shot voice cloning and speech synthesis capabilities, though it functions as a conversational assistant rather than a dedicated text-to-speech engine.

    PythonVoice CloningVoice Cloning EnginesZero-Shot Voice Cloning
    View on GitHub↗23,850
  • sesameailabs/csmSesameAILabs avatar

    SesameAILabs/csm

    14,669View on GitHub↗

    CSM is a conversational speech generation model and text-to-speech engine that converts text and audio inputs into synthetic speech. It utilizes a large language model architecture to predict and decode audio tokens for voice synthesis. The system functions as a zero-shot voice cloner, replicating specific speaker identities using short audio samples without requiring additional training. This enables precise control over speaker identity and the creation of synthetic speech that mimics a specific person. The model covers conversational speech synthesis and text-to-speech generation, transfo

    CSM is an open-source conversational speech generation model and text-to-speech engine capable of zero-shot voice cloning using audio samples, matching the core intent though it lacks explicit streaming audio or self-hosting infrastructure out of the box.

    PythonZero-Shot Voice CloningText-to-Speech Engines
    View on GitHub↗14,669
  • suno-ai/barksuno-ai avatar

    suno-ai/bark

    39,159View on GitHub↗

    Bark is a generative audio engine and machine learning inference library designed to convert written text into high-fidelity speech and sound effects. It functions as a text-to-audio transformer, utilizing multi-stage neural network architectures to map semantic input tokens into detailed audio codebooks for synthesis. The system distinguishes itself through a hierarchical transformer stacking approach that separates semantic understanding from acoustic realization. By employing autoregressive token prediction and vector quantized codebook mapping, the engine bridges linguistic and sonic doma

    Bark is a neural text-to-speech engine that generates realistic speech and non-speech sounds from text using transformer models, though it is primarily a Python library and research model rather than a complete ready-to-use voice cloning suite.

    Jupyter NotebookSpeech Synthesis ModelsText-to-Speech Engines
    View on GitHub↗39,159
  • microsoft/vibevoicemicrosoft avatar

    microsoft/VibeVoice

    49,394View on GitHub↗

    VibeVoice is a generative artificial intelligence platform designed for text-to-speech synthesis. It functions as a neural audio generation framework that converts written text into natural-sounding spoken audio, specifically engineered to maintain consistent vocal characteristics and narrative prosody across extended passages of content. The system distinguishes itself through its ability to generate long-form conversational speech while preserving speaker identity and linguistic content. By utilizing latent space disentanglement, the model separates speaker traits from the input text, allow

    VibeVoice is a neural text-to-speech platform focused on generative audio and speaker identity preservation, matching the core AI voice generation category and supporting voice cloning tools.

    PythonVoice Cloning Tools
    View on GitHub↗49,394
  • netease-youdao/emotivoicenetease-youdao avatar

    netease-youdao/EmotiVoice

    8,446View on GitHub↗

    EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio in English and Chinese. It utilizes a deep learning architecture to produce high-fidelity speech with controllable emotional states and timbres. The project includes a voice cloning framework for replicating specific speaker identities by training custom acoustic models on personal audio datasets. It employs a jointly-trained acoustic-vocoder pipeline and style-embedding-based synthesis to manage expression and reduce audio artifacts. The system covers a broad range of speec

    EmotiVoice is a neural text-to-speech engine that supports voice cloning, emotional synthesis, and self-hosting with a Python codebase, though it covers a narrower bilingual scope compared to fully multi-lingual engines.

    PythonVoice Cloning
    View on GitHub↗8,446
  • huggingface/speech-to-speechhuggingface avatar

    huggingface/speech-to-speech

    4,895View on GitHub↗

    This project is a framework for building local voice assistants and a real-time audio streaming server. It functions as a containerized inference engine and a multilingual speech pipeline that orchestrates speech-to-text, language models, and text-to-speech components to convert spoken input into spoken output. The system is distinguished by its use of WebSocket-based bidirectional streaming for low-latency interactions. It features a voice activity detection system that manages speech boundaries and handles user barge-in interruptions during assistant playback. It also supports custom voice

    This project is a real-time conversational speech pipeline and containerized inference engine that orchestrates speech components including text-to-speech and voice cloning, though its primary focus is on building voice assistants rather than standalone text-to-speech generation.

    PythonVoice Cloning Tools
    View on GitHub↗4,895
  • nvidia/nemoNVIDIA avatar

    NVIDIA/NeMo

    17,394View on GitHub↗

    NeMo is a multimodal AI framework and toolkit designed for the development, training, and scaling of large language models, generative AI systems, and speech-based models. It functions as an automatic speech recognition toolkit, a text-to-speech engine, and a framework for building models that process and generate combinations of text, image, and audio data. The project serves as a conversational AI orchestrator capable of managing real-time, interruptible voice interactions. It provides specialized workflows for speech translation, converting spoken audio from one language into text or speec

    NVIDIA NeMo is a comprehensive multimodal AI framework and toolkit that provides text-to-speech engines and voice processing capabilities, though it functions more as a broad model training and development framework rather than a ready-to-use plug-and-play voice cloning application.

    PythonAutomatic Speech RecognitionLarge Language Model Training FrameworksAutomatic Speech Recognition
    View on GitHub↗17,394
  • nvidia/tacotron2NVIDIA avatar

    NVIDIA/tacotron2

    5,300View on GitHub↗

    This project is a neural text-to-speech framework and PyTorch model designed to synthesize human speech. It converts written text into synthetic audio by predicting mel spectrograms, which serve as an intermediate representation for voice generation. The system includes a conditioning model for WaveNet to ensure natural-sounding audio output. It provides a distributed training framework that utilizes multi-GPU processing and automatic mixed precision to optimize training speed and reduce memory usage. The project covers the full pipeline of neural speech synthesis, from model training using

    This PyTorch implementation of Tacotron 2 provides a neural text-to-speech framework for synthesizing speech from text, though it lacks direct voice cloning capabilities and modern out-of-the-box pretrained models.

    Jupyter NotebookSpeech SynthesisConvolutional Encoder-DecodersInput Sequence Attentions
    View on GitHub↗5,300
Compare the top 10 at a glance
RepositoryStarsLanguageLicenseLast push
boson-ai/higgs-audio7.9KPythonapache-2.0Jan 18, 2026
babysor/mockingbird36.9KPythonNOASSERTIONMar 3, 2026
mozilla/tts10.2KJupyter NotebookMPL-2.0Nov 9, 2023
corentinj/real-time-voice-cloning59.9KPythonNOASSERTIONMar 9, 2026
myshell-ai/openvoice36.7KPythonMITApr 19, 2025
zyphra/zonos7.2KPythonApache-2.0Mar 5, 2025
coqui-ai/tts45.6KPythonMPL-2.0Aug 16, 2024
swivid/f5-tts14.8KPythonMITMay 18, 2026
funaudiollm/cosyvoice21.7KPythonApache-2.0May 25, 2026
koljab/realtimetts4KPythonMITMay 31, 2026

Related searches

  • an open source text to speech tool
  • an open source alternative to OpenAI API
  • an open source platform for LLM hosting
  • an open source real time voice changer
  • an open source llm proxy and gateway
  • an open source alternative to Otter.ai
  • an open source platform for AI characters
  • an open source alternative to HeyGen