awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
myshell-ai avatar

myshell-ai/OpenVoice

0
View on GitHub↗
36,720 stele·4,100 fork-uri·Python·MIT·15 vizualizăriresearch.myshell.ai/open-voice↗

OpenVoice

OpenVoice is a multilingual text-to-speech framework and voice cloning AI model designed for high-fidelity voice replication and low-latency audio generation. It functions as an instant speech synthesis engine that converts text to audio while replicating a specific speaker's tone and color.

The system is distinguished by its ability to perform cross-lingual cloning, allowing the vocal characteristics of a reference speaker to be applied to speech in different languages regardless of the original training data. It utilizes a decoupled representation to separate the physical identity of a voice from its emotional and rhythmic delivery.

This tool provides granular speech control over audio generation, enabling adjustments to parameters such as emotion, accent, rhythm, and intonation. These capabilities allow for the creation of digital replicas using short audio samples to synthesize expressive speech.

Features

  • Speech Synthesis Engines - Functions as a high-fidelity speech synthesis engine that converts text to audio with low latency.
  • Neural Text-to-Speech Engines - Implements a neural text-to-speech engine that combines text with style and tone vectors for audio generation.
  • Zero-Shot Voice Cloning - Extracts unique tone embeddings from short audio clips to replicate voices without needing additional model training.
  • Multilingual Synthesis - Synthesizes natural sounding speech across multiple languages while preserving a specific speaker's unique characteristics.
  • Text-to-Speech - Provides a multilingual framework for synthesizing natural human speech from text input.
  • Voice Cloning - Clones the specific tone color of a reference speaker to generate high-fidelity synthetic speech.
  • Cross-Lingual Voice Transfer - Provides the ability to transfer a speaker's vocal identity across different languages regardless of training data.
  • Prosody and Style Control - Allows granular adjustment of speech parameters such as rhythm, intonation, and emotion during the inference process.
  • Controllable Speech Generation - Provides a system for adjusting granular speech parameters such as emotion, accent, rhythm, and intonation.
  • Expressive Prosody Controls - Provides granular control over emotion, rhythm, and intonation to make digital voices sound more human.
  • Identity-Style Decoupling - Separates the physical identity of a voice from its emotional and rhythmic delivery to allow independent control.
  • Speech Processing - Instant voice cloning and speech synthesis.
  • Speech Synthesis - Instant voice cloning and speech synthesis.

Istoric stele

Graficul istoricului de stele pentru myshell-ai/openvoiceGraficul istoricului de stele pentru myshell-ai/openvoice

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Start searching with AI

Întrebări frecvente

Ce face myshell-ai/openvoice?

OpenVoice is a multilingual text-to-speech framework and voice cloning AI model designed for high-fidelity voice replication and low-latency audio generation. It functions as an instant speech synthesis engine that converts text to audio while replicating a specific speaker's tone and color.

Care sunt principalele funcționalități ale myshell-ai/openvoice?

Principalele funcționalități ale myshell-ai/openvoice sunt: Speech Synthesis Engines, Neural Text-to-Speech Engines, Zero-Shot Voice Cloning, Multilingual Synthesis, Text-to-Speech, Voice Cloning, Cross-Lingual Voice Transfer, Prosody and Style Control.

Care sunt câteva alternative open-source pentru myshell-ai/openvoice?

Alternativele open-source pentru myshell-ai/openvoice includ: boson-ai/higgs-audio — Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large… openbmb/voxcpm — VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice… hexgrad/kokoro — Kokoro is a lightweight neural text-to-speech engine that converts written text into spoken audio using a compact… fishaudio/fish-speech — This project is a generative speech synthesis engine that converts text into high-fidelity human speech. It utilizes a… nari-labs/dia — Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of… bytedance/megatts3 — MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English,…

Alternative open-source pentru OpenVoice

Proiecte open-source similare, clasificate după numărul de funcționalități comune cu OpenVoice.
  • boson-ai/higgs-audioAvatar boson-ai

    boson-ai/higgs-audio

    7,919Vezi pe GitHub↗

    Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large language model architectures. It functions as a multilingual speech synthesizer capable of generating high-fidelity audio across different languages with control over emotional tone and prosody. The system includes a voice cloning tool that creates synthetic replicas of specific speakers from short audio samples without requiring extensive model training. It also provides a streaming audio API designed to deliver generated speech incrementally to minimize playback delay. The

    Python
    Vezi pe GitHub↗7,919
  • openbmb/voxcpmAvatar OpenBMB

    OpenBMB/VoxCPM

    29,985Vezi pe GitHub↗

    VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice cloning tool and a synthetic voice designer, capable of generating natural speech across global languages and regional dialects using a GPU-accelerated audio generator. The project features a speech model fine-tuning framework that supports both full parameter updates and low-rank adaptation for customizing voice characteristics. It enables high-fidelity voice cloning from reference audio, including cross-lingual voice transfer and acoustic environment mimicry, as well as the crea

    Pythonaudiodeeplearningminicpm
    Vezi pe GitHub↗29,985
  • hexgrad/kokoroAvatar hexgrad

    hexgrad/kokoro

    5,729Vezi pe GitHub↗

    Kokoro is a lightweight neural text-to-speech engine that converts written text into spoken audio using a compact model designed for fast inference. It supports multiple languages through language-specific grapheme-to-phoneme conversion pipelines, and offers voice profile selection to change the character of the generated speech. The engine provides GPU acceleration on Apple Silicon hardware by setting a single environment variable, enabling faster inference on Mac M-series machines. It also includes pattern-based text segmentation, allowing input text to be split at user-defined delimiters t

    JavaScript
    Vezi pe GitHub↗5,729
  • fishaudio/fish-speechAvatar fishaudio

    fishaudio/fish-speech

    24,928Vezi pe GitHub↗

    This project is a generative speech synthesis engine that converts text into high-fidelity human speech. It utilizes a two-stage autoregressive transformer architecture that separates semantic token prediction from acoustic detail reconstruction to balance linguistic accuracy with audio quality. The system is designed to support multilingual output and conversational AI development, enabling the generation of context-aware speech that maintains flow across multiple dialogue turns. The platform distinguishes itself through a production-ready inference server that employs continuous batching to

    Pythonllamatransformertts
    Vezi pe GitHub↗24,928
Vezi toate cele 30 alternative pentru OpenVoice→