awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
jasonppy avatar

jasonppy/VoiceCraft

0
View on GitHub↗
8,500 estrellas·796 forks·Jupyter Notebook·13 vistas

VoiceCraft

VoiceCraft is a neural speech generation and manipulation system consisting of a text-to-speech system, a voice cloning tool, and an audio inpainting engine. It uses a large language model approach to synthesize high-fidelity audio from text and replicate speaker identities.

The system provides zero-shot voice cloning and speech editing capabilities, allowing users to modify spoken content within existing recordings. This includes an audio inpainting engine that replaces specific sections of audio with new speech while preserving the original acoustic characteristics and speaker identity.

The project covers high-level capabilities for text-to-speech synthesis, custom voice model training through phoneme-based tokenization, and acoustic speech refinement. It utilizes autoregressive synthesis and latent space representations to decouple speaker identity from linguistic content.

Features

  • Zero-Shot Voice Cloning - Generates high-fidelity speech using short reference audio samples to replicate speaker identity without retraining.
  • Voice Cloning Tools - Provides a pipeline to generate high-quality synthetic speech by processing custom audio recordings and transcripts.
  • Voice Model Trainers - Converts audio recordings and transcripts into phoneme sequences to train and refine neural speech models.
  • Phoneme-Based Pipelines - Converts text and audio transcripts into discrete phonetic units to standardize speech generation.
  • Audio Inpainting And Editing - Provides tools for modifying and regenerating specific segments of existing audio using text-based guidance.
  • Autoregressive Synthesis - Implements autoregressive audio synthesis to produce natural speech rhythms and prosody from text input.
  • Text-to-Speech - Synthesizes natural human speech from text input using high-fidelity neural generative models.
  • Surgical Audio Editing - Allows for the modification of spoken content within existing recordings while preserving original voice identity.
  • Voice Cloning - Replicates specific human vocal characteristics from audio samples to create high-fidelity synthetic voice models.
  • Audio Content Refinement - Replacing or correcting specific words in a recording without needing to re-record the entire session.
  • Audio Gap Infilling - Provides a neural engine for predicting and restoring missing audio segments to modify spoken content.
  • Latent Space Encoders - Utilizes latent space encoders to decouple speaker identity from linguistic content for synthetic generation.
  • Acoustic Models - Uses neural acoustic models to convert linguistic representations into high-fidelity audio features.
  • Zero-Shot Speech Editors - Modifies spoken content and infills audio tokens while preserving original voice identity without retraining.
  • Speech-to-Speech with Video Streams - Replaces sections of existing audio with new speech while maintaining original acoustic characteristics.

Historial de estrellas

Gráfico del historial de estrellas de jasonppy/voicecraftGráfico del historial de estrellas de jasonppy/voicecraft

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Preguntas frecuentes

¿Qué hace jasonppy/voicecraft?

VoiceCraft is a neural speech generation and manipulation system consisting of a text-to-speech system, a voice cloning tool, and an audio inpainting engine. It uses a large language model approach to synthesize high-fidelity audio from text and replicate speaker identities.

¿Cuáles son las características principales de jasonppy/voicecraft?

Las características principales de jasonppy/voicecraft son: Zero-Shot Voice Cloning, Voice Cloning Tools, Voice Model Trainers, Phoneme-Based Pipelines, Audio Inpainting And Editing, Autoregressive Synthesis, Text-to-Speech, Surgical Audio Editing.

¿Qué alternativas de código abierto existen para jasonppy/voicecraft?

Las alternativas de código abierto para jasonppy/voicecraft incluyen: babysor/mockingbird — MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions… netease-youdao/emotivoice — EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio… elevenlabs/elevenlabs-python — This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of… bytedance/megatts3 — MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English,… getstream/vision-agents. swivid/f5-tts — F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent…

Alternativas open-source a VoiceCraft

Proyectos open-source similares, clasificados según cuántas características comparten con VoiceCraft.
  • babysor/mockingbirdAvatar de babysor

    babysor/MockingBird

    36,903Ver en GitHub↗

    MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions as a voice synthesis trainer for building custom models from audio datasets, a command-line generator for producing audio files, and a text-to-speech server for remote application integration. The project specializes in real-time voice cloning, which extracts vocal characteristics from short audio samples to mimic a target speaker's unique timbre. It utilizes reference-driven audio synthesis to condition pre-trained models on specific audio samples, allowing for the generation

    Pythonaideep-learningpytorch
    Ver en GitHub↗36,903
  • elevenlabs/elevenlabs-pythonAvatar de elevenlabs

    elevenlabs/elevenlabs-python

    2,873Ver en GitHub↗

    This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of conversational AI agents. It enables the creation of lifelike spoken audio from text, the replication of human voices through custom cloning, and the deployment of real-time voice agents capable of interacting with external large language models. The library distinguishes itself through deep integration of conversational AI capabilities, including the design of agent personas and the execution of real-time actions via APIs. It supports professional-grade audio production thro

    Pythonartificial-intelligenceconversational-aitext-to-speech
    Ver en GitHub↗2,873
  • bytedance/megatts3Avatar de bytedance

    bytedance/MegaTTS3

    6,066Ver en GitHub↗

    MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English, including seamless code-switching within a single utterance. It functions as a text-to-speech engine, voice cloning system, and speech-to-text alignment tool, built around an acoustic latent compression model that encodes high-resolution audio into compact representations for efficient processing. The system distinguishes itself through accent intensity control, allowing adjustment of a speaker's accent strength in generated speech, and voice cloning from short audio samples for pers

    Pythonresearch
    Ver en GitHub↗6,066
  • netease-youdao/emotivoiceAvatar de netease-youdao

    netease-youdao/EmotiVoice

    8,446Ver en GitHub↗

    EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio in English and Chinese. It utilizes a deep learning architecture to produce high-fidelity speech with controllable emotional states and timbres. The project includes a voice cloning framework for replicating specific speaker identities by training custom acoustic models on personal audio datasets. It employs a jointly-trained acoustic-vocoder pipeline and style-embedding-based synthesis to manage expression and reduce audio artifacts. The system covers a broad range of speec

    Pythonaideep-learningemotion
    Ver en GitHub↗8,446
  • Ver las 30 alternativas a VoiceCraft→