awesome-repositories.comCatégoriesBlog
MCP
awesome-repositories.com

Découvrez les meilleurs dépôts open-source grâce à notre recherche par IA.

ExplorerRecherches sélectionnéesAlternatives open sourceLogiciels auto-hébergésBlogPlan du site
ProjetServeur MCPÀ proposNotre méthodologiePresse
Mentions légalesConfidentialitéConditions d'utilisation
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
FunAudioLLM avatar

FunAudioLLM/CosyVoice

0
View on GitHub↗
21,673 stars·2,498 forks·Python·Apache-2.0·31 vuesfunaudiollm.github.io/cosyvoice3↗

CosyVoice

CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones.

The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjust prosody, pacing, volume, and breathing to achieve realistic output. Furthermore, the system supports phoneme-level alignment and latent space conditioning to modulate emotional personas and ensure precise pronunciation.

The architecture incorporates reinforcement learning to iteratively refine output quality and alignment with human-perceived speech standards. Users can also perform custom speaker model adaptation to improve voice similarity and consistency for specialized production requirements.

Features

  • Neural Text-to-Speech Engines - Functions as a speech synthesis framework using large language models to generate expressive, multilingual audio.
  • Zero-Shot Voice Cloning - Enables the creation of synthetic voice profiles from short audio samples without requiring additional model training.
  • Speech Synthesis - Generates natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones.
  • Expressive Synthesis Models - Implements a neural synthesis architecture that modulates vocal attributes to produce speech with customizable emotional personas.
  • Autoregressive Models - Generates speech by predicting sequences of discrete acoustic tokens using a transformer architecture.
  • Prosody Control Tokens - Injects emotional and stylistic vector representations into the model to modulate prosody and tone.
  • Multilingual Speech Models - Generates natural-sounding audio supporting multiple languages and mixed-lingual content.
  • Prosody Controls - Provides fine-grained control over vocal attributes like breathing, pacing, and volume to produce realistic, human-like speech output.
  • Cross-Modal Alignment Models - Conditions speech generation on both text input and reference audio embeddings to align synthetic output with target speaker characteristics.
  • Reinforcement Learning Alignment - Refines model outputs using reward signals to optimize for naturalness and human-perceived speech quality.
  • Phoneme-Based Alignment - Maps text inputs to specific phonetic sequences to ensure precise pronunciation and prosodic rendering.
  • Speech Processing - Multilingual speech generation and cloning model.
  • Speech Synthesis - Multilingual speech synthesis and voice cloning.
  • Text to speech - Listed in the “Text to speech” section of the Ailia Models awesome list.
  • Emotional Modulation - Generates multilingual speech while applying specific emotional tones for engaging communication.
  • Neural Vocoders - Converts raw acoustic tokens into high-fidelity waveforms using deep learning models.
  • Custom Model Adapters - Refines base speech generation models for specific target speakers to improve voice similarity and consistency.
  • Prosody Modulation Tools - Provides controls to modify the emotional tone and prosodic style of generated audio.
  • Speech Model Fine-Tuning - Refines base speech generation models for specific target speakers to improve voice similarity.
  • Dialectal Synthesis Engines - Generates audio while preserving specific phonetic and tonal characteristics of regional dialects.
  • Phonetic Pronunciation Overrides - Allows overriding default speech output with explicit phoneme sequences for precise pronunciation.

Historique des stars

Graphique de l'historique des stars pour funaudiollm/cosyvoiceGraphique de l'historique des stars pour funaudiollm/cosyvoice

Recherche par IA

Explorez plus de dépôts awesome

Décrivez vos besoins en langage naturel — l'IA classe des milliers de projets open source sélectionnés par pertinence.

Start searching with AI

Alternatives open source à CosyVoice

Projets open source similaires, classés selon le nombre de fonctionnalités partagées avec CosyVoice.
  • openbmb/voxcpmAvatar de OpenBMB

    OpenBMB/VoxCPM

    29,985Voir sur GitHub↗

    VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice cloning tool and a synthetic voice designer, capable of generating natural speech across global languages and regional dialects using a GPU-accelerated audio generator. The project features a speech model fine-tuning framework that supports both full parameter updates and low-rank adaptation for customizing voice characteristics. It enables high-fidelity voice cloning from reference audio, including cross-lingual voice transfer and acoustic environment mimicry, as well as the crea

    Pythonaudiodeeplearningminicpm
    Voir sur GitHub↗29,985
  • nari-labs/diaAvatar de nari-labs

    nari-labs/dia

    19,324Voir sur GitHub↗

    Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt

    Pythonaiopen-weighttext-to-speech
    Voir sur GitHub↗19,324
  • boson-ai/higgs-audioAvatar de boson-ai

    boson-ai/higgs-audio

    7,919Voir sur GitHub↗

    Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large language model architectures. It functions as a multilingual speech synthesizer capable of generating high-fidelity audio across different languages with control over emotional tone and prosody. The system includes a voice cloning tool that creates synthetic replicas of specific speakers from short audio samples without requiring extensive model training. It also provides a streaming audio API designed to deliver generated speech incrementally to minimize playback delay. The

    Python
    Voir sur GitHub↗7,919
  • sparkaudio/spark-ttsAvatar de SparkAudio

    SparkAudio/Spark-TTS

    10,930Voir sur GitHub↗

    Spark-TTS is a deep learning text-to-speech synthesis engine designed to convert written text into high-fidelity audio. It utilizes a transformer-based architecture and autoregressive sequence modeling to generate coherent speech, transforming linguistic input into natural-sounding waveforms through neural speech codec synthesis. The platform distinguishes itself through zero-shot voice cloning, which allows users to mimic a target speaker’s unique vocal identity using only a short reference audio sample without requiring additional model training. It also features cross-lingual phonetic mapp

    Python
    Voir sur GitHub↗10,930
Voir les 30 alternatives à CosyVoice→

Questions fréquentes

Que fait funaudiollm/cosyvoice ?

CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones.

Quelles sont les fonctionnalités principales de funaudiollm/cosyvoice ?

Les fonctionnalités principales de funaudiollm/cosyvoice sont : Neural Text-to-Speech Engines, Zero-Shot Voice Cloning, Speech Synthesis, Expressive Synthesis Models, Autoregressive Models, Prosody Control Tokens, Multilingual Speech Models, Prosody Controls.

Quelles sont les alternatives open-source à funaudiollm/cosyvoice ?

Les alternatives open-source à funaudiollm/cosyvoice incluent : openbmb/voxcpm — VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice… nari-labs/dia — Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of… boson-ai/higgs-audio — Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large… sparkaudio/spark-tts — Spark-TTS is a deep learning text-to-speech synthesis engine designed to convert written text into high-fidelity… index-tts/index-tts — Index-tts is a neural audio generation engine designed to convert written text into high-fidelity human speech. By… rvc-boss/gpt-sovits — GPT-SoVITS is a text-to-speech synthesis engine and voice cloning toolkit designed for generating natural-sounding…