awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
WhisperSpeech avatar

WhisperSpeech/WhisperSpeech

0
View on GitHub↗
4,617 stars·271 forks·Jupyter Notebook·MIT·28 viewswhisperspeech.github.io/WhisperSpeech↗

WhisperSpeech

WhisperSpeech is a multilingual speech synthesizer and neural text-to-speech system. It functions by inverting the Whisper model architecture to convert text into high-fidelity synthetic audio.

The system enables voice cloning by using reference audio files to mimic specific speakers. It supports multilingual speech production, which includes the ability to generate audio across different languages and handle language switching within a single sentence.

The project covers a broad range of speech capabilities, including text-to-speech generation and speech dataset preparation. It incorporates tools for transcribing speech to text, extracting acoustic tokens, and detecting voice activity.

Features

  • Text-to-Speech - Generates high-fidelity synthetic audio from text using a neural multi-stage token pipeline.
  • Voice Cloning Tools - Enables mimicking specific speakers using reference audio files to guide synthetic speech generation.
  • Zero-Shot Voice Cloning - Provides zero-shot voice cloning by extracting speaker embeddings from short reference audio clips.
  • Multilingual Text-to-Speech Engines - Specializes in multilingual speech production with seamless mixing of multiple languages in one output.
  • Voice Cloning Engines - Mimics specific human voices by using reference audio samples to guide the synthesis engine.
  • Inverted Architecture Models - Utilizes an inverted Whisper model architecture to convert text into high-fidelity synthetic audio.
  • Multilingual Speech Synthesizers - Produces spoken audio in multiple languages with the ability to switch languages within a single sentence.
  • Cross-Lingual Semantic Mappings - Implements a shared semantic space to enable seamless language switching within a single synthetic speech stream.
  • Acoustic Token Pipelines - Ships a multi-stage pipeline that separates linguistic and sonic features via semantic and acoustic tokenization.
  • Speech Dataset Engineering - Provides a complete pipeline for transcribing audio and extracting tokens to build speech synthesis training sets.
  • Speech-to-Text Transcribers - Transcribes audio segments into text and semantic tokens to create high-quality training datasets.
  • Neural Audio Pipelines - Implements an end-to-end neural pipeline using semantic and acoustic tokens to generate high-fidelity synthetic speech.
  • Discrete Token Extraction - Uses discrete token extraction to represent audio as quantized sequences for language-model-based speech generation.

Star history

Star history chart for whisperspeech/whisperspeechStar history chart for whisperspeech/whisperspeech

How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to WhisperSpeech

Similar open-source projects, ranked by how many features they share with WhisperSpeech.
  • zyphra/zonosZyphra avatar

    Zyphra/Zonos

    7,225View on GitHub↗

    Zonos is a controllable audio synthesis engine and large language model for text-to-speech. It serves as a multilingual speech generator capable of producing audio in English, Japanese, Chinese, French, and German. The system provides zero-shot voice cloning, allowing the replication of specific human voices using short audio samples. It supports the capture of nuanced behaviors, such as whispering, and provides parametric control over speaking rate, pitch, frequency, and emotional tone. The project covers a broad range of expressive speech synthesis and custom audio generation capabilities,

    Python
    View on GitHub↗7,225
  • kevinwang676/bark-voice-cloningKevinWang676 avatar

    KevinWang676/Bark-Voice-Cloning

    2,957View on GitHub↗

    Bark Voice Cloning is a text-to-speech synthesis engine designed to generate natural-sounding audio and replicate specific vocal characteristics. The system utilizes a transformer-based autoregressive model to convert written text into high-fidelity speech, supporting multilingual output and expressive delivery. The project distinguishes itself through zero-shot voice cloning, which extracts speaker identity embeddings from short audio samples to condition the generative model without requiring extensive fine-tuning. It also provides specialized workflows for voice identity conversion, allowi

    Jupyter Notebook
    View on GitHub↗2,957
  • babysor/mockingbirdbabysor avatar

    babysor/MockingBird

    36,903View on GitHub↗

    MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions as a voice synthesis trainer for building custom models from audio datasets, a command-line generator for producing audio files, and a text-to-speech server for remote application integration. The project specializes in real-time voice cloning, which extracts vocal characteristics from short audio samples to mimic a target speaker's unique timbre. It utilizes reference-driven audio synthesis to condition pre-trained models on specific audio samples, allowing for the generation

    Pythonaideep-learningpytorch
    View on GitHub↗36,903
  • jasonppy/voicecraftjasonppy avatar

    jasonppy/VoiceCraft

    8,500View on GitHub↗

    VoiceCraft is a neural speech generation and manipulation system consisting of a text-to-speech system, a voice cloning tool, and an audio inpainting engine. It uses a large language model approach to synthesize high-fidelity audio from text and replicate speaker identities. The system provides zero-shot voice cloning and speech editing capabilities, allowing users to modify spoken content within existing recordings. This includes an audio inpainting engine that replaces specific sections of audio with new speech while preserving the original acoustic characteristics and speaker identity. Th

    Jupyter Notebook
    View on GitHub↗8,500
See all 30 alternatives to WhisperSpeech→

Frequently asked questions

What does whisperspeech/whisperspeech do?

WhisperSpeech is a multilingual speech synthesizer and neural text-to-speech system. It functions by inverting the Whisper model architecture to convert text into high-fidelity synthetic audio.

What are the main features of whisperspeech/whisperspeech?

The main features of whisperspeech/whisperspeech are: Text-to-Speech, Voice Cloning Tools, Zero-Shot Voice Cloning, Multilingual Text-to-Speech Engines, Voice Cloning Engines, Inverted Architecture Models, Multilingual Speech Synthesizers, Cross-Lingual Semantic Mappings.

What are some open-source alternatives to whisperspeech/whisperspeech?

Open-source alternatives to whisperspeech/whisperspeech include: zyphra/zonos — Zonos is a controllable audio synthesis engine and large language model for text-to-speech. It serves as a… kevinwang676/bark-voice-cloning — Bark Voice Cloning is a text-to-speech synthesis engine designed to generate natural-sounding audio and replicate… swivid/f5-tts — F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent… babysor/mockingbird — MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions… jasonppy/voicecraft — VoiceCraft is a neural speech generation and manipulation system consisting of a text-to-speech system, a voice… metavoiceio/metavoice-src — This project is an expressive text-to-speech foundation model and voice cloning system designed to synthesize…