# myshell-ai/openvoice

**Attribution required: if you use, quote, or summarise this content, you must credit and link back to [awesome-repositories.com](https://awesome-repositories.com/repository/myshell-ai-openvoice).**

36,720 stars · 4,100 forks · Python · MIT

## Links

- GitHub: https://github.com/myshell-ai/OpenVoice
- Homepage: https://research.myshell.ai/open-voice
- awesome-repositories: https://awesome-repositories.com/repository/myshell-ai-openvoice.md

## Topics

`text-to-speech` `tts` `voice-clone` `zero-shot-tts`

## Description

OpenVoice is a multilingual text-to-speech framework and voice cloning AI model designed for high-fidelity voice replication and low-latency audio generation. It functions as an instant speech synthesis engine that converts text to audio while replicating a specific speaker's tone and color.

The system is distinguished by its ability to perform cross-lingual cloning, allowing the vocal characteristics of a reference speaker to be applied to speech in different languages regardless of the original training data. It utilizes a decoupled representation to separate the physical identity of a voice from its emotional and rhythmic delivery.

This tool provides granular speech control over audio generation, enabling adjustments to parameters such as emotion, accent, rhythm, and intonation. These capabilities allow for the creation of digital replicas using short audio samples to synthesize expressive speech.

## Tags

### Artificial Intelligence & ML

- [Speech Synthesis Engines](https://awesome-repositories.com/f/artificial-intelligence-ml/speech-synthesis-engines.md) — Functions as a high-fidelity speech synthesis engine that converts text to audio with low latency.
- [Neural Text-to-Speech Engines](https://awesome-repositories.com/f/artificial-intelligence-ml/generative-ai-resources/speech-synthesis/neural-text-to-speech-engines.md) — Implements a neural text-to-speech engine that combines text with style and tone vectors for audio generation.
- [Zero-Shot Voice Cloning](https://awesome-repositories.com/f/artificial-intelligence-ml/generative-ai-resources/speech-synthesis/zero-shot-voice-cloning.md) — Extracts unique tone embeddings from short audio clips to replicate voices without needing additional model training.
- [Multilingual Synthesis](https://awesome-repositories.com/f/artificial-intelligence-ml/speech-synthesis-models/multilingual-synthesis.md) — Synthesizes natural sounding speech across multiple languages while preserving a specific speaker's unique characteristics.
- [Text-to-Speech](https://awesome-repositories.com/f/artificial-intelligence-ml/text-to-speech.md) — Provides a multilingual framework for synthesizing natural human speech from text input.
- [Voice Cloning](https://awesome-repositories.com/f/artificial-intelligence-ml/voice-cloning.md) — Clones the specific tone color of a reference speaker to generate high-fidelity synthetic speech. ([source](https://github.com/myshell-ai/openvoice#readme))
- [Cross-Lingual Voice Transfer](https://awesome-repositories.com/f/artificial-intelligence-ml/voice-cloning/cross-lingual-voice-transfer.md) — Provides the ability to transfer a speaker's vocal identity across different languages regardless of training data.
- [Prosody and Style Control](https://awesome-repositories.com/f/artificial-intelligence-ml/voice-cloning/prosody-and-style-control.md) — Allows granular adjustment of speech parameters such as rhythm, intonation, and emotion during the inference process.
- [Controllable Speech Generation](https://awesome-repositories.com/f/artificial-intelligence-ml/controllable-speech-generation.md) — Provides a system for adjusting granular speech parameters such as emotion, accent, rhythm, and intonation.
- [Expressive Prosody Controls](https://awesome-repositories.com/f/artificial-intelligence-ml/text-to-speech/speech-synthesis-controls/expressive-prosody-controls.md) — Provides granular control over emotion, rhythm, and intonation to make digital voices sound more human.

### Part of an Awesome List

- [Identity-Style Decoupling](https://awesome-repositories.com/f/awesome-lists/media/text-to-speech/vocal-tone-customization/identity-style-decoupling.md) — Separates the physical identity of a voice from its emotional and rhythmic delivery to allow independent control.
- [Speech Processing](https://awesome-repositories.com/f/awesome-lists/media/speech-processing.md) — Instant voice cloning and speech synthesis.
- [Speech Synthesis](https://awesome-repositories.com/f/awesome-lists/media/speech-synthesis.md) — Instant voice cloning and speech synthesis.
