awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
jasonppy avatar

jasonppy/VoiceCraft

0
View on GitHub↗
8,500 stars·796 forks·Jupyter Notebook·29 views

VoiceCraft

VoiceCraft is a neural speech generation and manipulation system consisting of a text-to-speech system, a voice cloning tool, and an audio inpainting engine. It uses a large language model approach to synthesize high-fidelity audio from text and replicate speaker identities.

The system provides zero-shot voice cloning and speech editing capabilities, allowing users to modify spoken content within existing recordings. This includes an audio inpainting engine that replaces specific sections of audio with new speech while preserving the original acoustic characteristics and speaker identity.

The project covers high-level capabilities for text-to-speech synthesis, custom voice model training through phoneme-based tokenization, and acoustic speech refinement. It utilizes autoregressive synthesis and latent space representations to decouple speaker identity from linguistic content.

Features

  • Zero-Shot Voice Cloning - Generates high-fidelity speech using short reference audio samples to replicate speaker identity without retraining.
  • Voice Cloning Tools - Provides a pipeline to generate high-quality synthetic speech by processing custom audio recordings and transcripts.
  • Voice Model Trainers - Converts audio recordings and transcripts into phoneme sequences to train and refine neural speech models.
  • Phoneme-Based Pipelines - Converts text and audio transcripts into discrete phonetic units to standardize speech generation.
  • Audio Inpainting And Editing - Provides tools for modifying and regenerating specific segments of existing audio using text-based guidance.
  • Autoregressive Synthesis - Implements autoregressive audio synthesis to produce natural speech rhythms and prosody from text input.
  • Text-to-Speech - Synthesizes natural human speech from text input using high-fidelity neural generative models.
  • Surgical Audio Editing - Allows for the modification of spoken content within existing recordings while preserving original voice identity.
  • Voice Cloning - Replicates specific human vocal characteristics from audio samples to create high-fidelity synthetic voice models.
  • Audio Content Refinement - Replacing or correcting specific words in a recording without needing to re-record the entire session.
  • Audio Gap Infilling - Provides a neural engine for predicting and restoring missing audio segments to modify spoken content.
  • Latent Space Encoders - Utilizes latent space encoders to decouple speaker identity from linguistic content for synthetic generation.
  • Acoustic Models - Uses neural acoustic models to convert linguistic representations into high-fidelity audio features.
  • Zero-Shot Speech Editors - Modifies spoken content and infills audio tokens while preserving original voice identity without retraining.
  • Speech-to-Speech with Video Streams - Replaces sections of existing audio with new speech while maintaining original acoustic characteristics.

Star history

Star history chart for jasonppy/voicecraftStar history chart for jasonppy/voicecraft

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with VoiceCraft

These projects share indexed features with VoiceCraft. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • babysor/mockingbirdbabysor avatar

    babysor/MockingBird

    36,903View on GitHub↗

    MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions as a voice synthesis trainer for building custom models from audio datasets, a command-line generator for producing audio files, and a text-to-speech server for remote application integration. The project specializes in real-time voice cloning, which extracts vocal characteristics from short audio samples to mimic a target speaker's unique timbre. It utilizes reference-driven audio synthesis to condition pre-trained models on specific audio samples, allowing for the generation

    Pythonaideep-learningpytorch
    View on GitHub↗36,903
  • elevenlabs/elevenlabs-pythonelevenlabs avatar

    elevenlabs/elevenlabs-python

    2,873View on GitHub↗

    This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of conversational AI agents. It enables the creation of lifelike spoken audio from text, the replication of human voices through custom cloning, and the deployment of real-time voice agents capable of interacting with external large language models. The library distinguishes itself through deep integration of conversational AI capabilities, including the design of agent personas and the execution of real-time actions via APIs. It supports professional-grade audio production thro

    Pythonartificial-intelligenceconversational-aitext-to-speech
    View on GitHub↗2,873
  • bytedance/megatts3bytedance avatar

    bytedance/MegaTTS3

    6,066View on GitHub↗

    MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English, including seamless code-switching within a single utterance. It functions as a text-to-speech engine, voice cloning system, and speech-to-text alignment tool, built around an acoustic latent compression model that encodes high-resolution audio into compact representations for efficient processing. The system distinguishes itself through accent intensity control, allowing adjustment of a speaker's accent strength in generated speech, and voice cloning from short audio samples for pers

    Pythonresearch
    View on GitHub↗6,066
  • netease-youdao/emotivoicenetease-youdao avatar

    netease-youdao/EmotiVoice

    8,446View on GitHub↗

    EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio in English and Chinese. It utilizes a deep learning architecture to produce high-fidelity speech with controllable emotional states and timbres. The project includes a voice cloning framework for replicating specific speaker identities by training custom acoustic models on personal audio datasets. It employs a jointly-trained acoustic-vocoder pipeline and style-embedding-based synthesis to manage expression and reduce audio artifacts. The system covers a broad range of speec

    Pythonaideep-learningemotion
    View on GitHub↗8,446
Compare all 30 related projects→

Frequently asked questions

What does jasonppy/voicecraft do?

VoiceCraft is a neural speech generation and manipulation system consisting of a text-to-speech system, a voice cloning tool, and an audio inpainting engine. It uses a large language model approach to synthesize high-fidelity audio from text and replicate speaker identities.

What are the main features of jasonppy/voicecraft?

The main features of jasonppy/voicecraft are: Zero-Shot Voice Cloning, Voice Cloning Tools, Voice Model Trainers, Phoneme-Based Pipelines, Audio Inpainting And Editing, Autoregressive Synthesis, Text-to-Speech, Surgical Audio Editing.

Which projects share features with jasonppy/voicecraft?

Projects with overlapping indexed features include: babysor/mockingbird — MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions… netease-youdao/emotivoice — EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio… elevenlabs/elevenlabs-python — This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of… bytedance/megatts3 — MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English,… getstream/vision-agents. swivid/f5-tts — F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent…