awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Plachtaa avatar

Plachtaa/seed-vcArchived

0
View on GitHub↗
3,590 stars·443 forks·Python·gpl-3.0·19 views

Seed Vc

seed-vc is an AI voice conversion tool and voice cloning system designed to transform the timbre, accent, and emotion of speech recordings. It provides a framework for replicating specific speaker identities and singing styles using short reference audio samples.

The project includes a voice fine-tuning framework for training models on custom audio datasets to increase the accuracy of voice clones. It also features speech anonymization tools that remove unique speaker traits to produce a generic average voice for identity protection.

The system covers a broad range of audio processing capabilities, including zero-shot voice conversion, talking pace control, and the modification of emotional delivery and accents. It supports both spoken speech and singing voice conversion to transfer styles between source and target recordings.

Features

  • Zero-Shot Voice Cloning - Transforms source speech into a target speaker identity from short samples without requiring model retraining.
  • Custom Model Training - Fine-tunes machine learning models on specialized audio datasets to increase the likeness of cloned speakers.
  • Fine-Tuning Frameworks - Provides a framework for training models on custom audio datasets to improve the accuracy of voice clones.
  • Voice Model Trainers - Trains and fine-tunes voice models on custom audio datasets to increase speaker similarity.
  • Voice Cloning - Creates high-fidelity digital replicas of specific people's voices using short audio samples.
  • Voice Identity Conversions - Clones speaker voice timbre using reference audio samples to transform source recordings.
  • Model Fine-Tuning - Provides a framework for optimizing pre-trained voice models on specific audio datasets to improve cloning accuracy.
  • Singing Style Transfer - Clones target singers' voices and singing styles using short reference samples to transform source recordings.
  • Vocal Timbre Extraction - Extracts vocal characteristics from short reference samples to reshape the spectral envelope of source recordings.
  • Neural Pace Control - Adjusts the temporal duration of speech waveforms to change talking pace while preserving audio quality.
  • Emotional Synthesis - Utilizes latent-space embeddings to incorporate controllable emotional states and accents into synthesized speech.
  • Speech Style Transfer - Changes the accent, emotion, and delivery of recordings while preserving or altering the original voice.
  • Vocal Identity Anonymization - Implements averaging-based identity anonymization to protect speaker privacy by producing a generic vocal profile.
  • Emotion and Accent Transformation - Modifies the accent and emotional delivery of source recordings while maintaining or altering speaker timbre.
  • Audio Identity Anonymization - Removes identifying vocal characteristics from recordings to produce a generic voice for identity protection.
  • Speech Anonymization - Provides tools that remove unique speaker traits from recordings to produce a generic average voice for identity protection.
  • Speech Anonymization Tools - Removes individual speaker traits from audio recordings to produce a generic average voice.

Star history

Star history chart for plachtaa/seed-vcStar history chart for plachtaa/seed-vc

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Seed Vc

These projects share indexed features with Seed Vc. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • babysor/mockingbirdbabysor avatar

    babysor/MockingBird

    36,903View on GitHub↗

    MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions as a voice synthesis trainer for building custom models from audio datasets, a command-line generator for producing audio files, and a text-to-speech server for remote application integration. The project specializes in real-time voice cloning, which extracts vocal characteristics from short audio samples to mimic a target speaker's unique timbre. It utilizes reference-driven audio synthesis to condition pre-trained models on specific audio samples, allowing for the generation

    Pythonaideep-learningpytorch
    View on GitHub↗36,903
  • jasonppy/voicecraftjasonppy avatar

    jasonppy/VoiceCraft

    8,500View on GitHub↗

    VoiceCraft is a neural speech generation and manipulation system consisting of a text-to-speech system, a voice cloning tool, and an audio inpainting engine. It uses a large language model approach to synthesize high-fidelity audio from text and replicate speaker identities. The system provides zero-shot voice cloning and speech editing capabilities, allowing users to modify spoken content within existing recordings. This includes an audio inpainting engine that replaces specific sections of audio with new speech while preserving the original acoustic characteristics and speaker identity. Th

    Jupyter Notebook
    View on GitHub↗8,500
  • plachtaa/vall-e-xPlachtaa avatar

    Plachtaa/VALL-E-X

    7,939View on GitHub↗

    VALL-E-X is a neural speech synthesis framework and zero-shot text-to-speech engine. It functions as a multilingual synthesizer capable of generating natural human speech with control over emotion, pitch, and prosody. The project specializes in zero-shot voice cloning and cross-lingual voice replication, allowing the system to produce personalized speech in multiple target languages using short audio samples without additional training. It further enables cross-language accent manipulation and the ability to match the emotional tone and acoustic environment of a provided prompt. The implemen

    Pythonemotional-speechgpttext-to-speech
    View on GitHub↗7,939
  • netease-youdao/emotivoicenetease-youdao avatar

    netease-youdao/EmotiVoice

    8,446View on GitHub↗

    EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio in English and Chinese. It utilizes a deep learning architecture to produce high-fidelity speech with controllable emotional states and timbres. The project includes a voice cloning framework for replicating specific speaker identities by training custom acoustic models on personal audio datasets. It employs a jointly-trained acoustic-vocoder pipeline and style-embedding-based synthesis to manage expression and reduce audio artifacts. The system covers a broad range of speec

    Pythonaideep-learningemotion
    View on GitHub↗8,446
Compare all 30 related projects→

Frequently asked questions

What does plachtaa/seed-vc do?

seed-vc is an AI voice conversion tool and voice cloning system designed to transform the timbre, accent, and emotion of speech recordings. It provides a framework for replicating specific speaker identities and singing styles using short reference audio samples.

What are the main features of plachtaa/seed-vc?

The main features of plachtaa/seed-vc are: Zero-Shot Voice Cloning, Custom Model Training, Fine-Tuning Frameworks, Voice Model Trainers, Voice Cloning, Voice Identity Conversions, Model Fine-Tuning, Singing Style Transfer.

Which projects share features with plachtaa/seed-vc?

Projects with overlapping indexed features include: babysor/mockingbird — MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions… jasonppy/voicecraft — VoiceCraft is a neural speech generation and manipulation system consisting of a text-to-speech system, a voice… plachtaa/vall-e-x — VALL-E-X is a neural speech synthesis framework and zero-shot text-to-speech engine. It functions as a multilingual… netease-youdao/emotivoice — EmotiVoice is an emotional text-to-speech engine and bilingual speech synthesizer designed to generate synthetic audio… kevinwang676/bark-voice-cloning — Bark Voice Cloning is a text-to-speech synthesis engine designed to generate natural-sounding audio and replicate… coqui-ai/tts — This project is a deep learning text-to-speech toolkit used for training and deploying neural speech synthesis models.…