awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
2noise avatar

2noise/ChatTTS

0
View on GitHub↗
39,464 estrellas·4,246 forks·Python·AGPL-3.0·18 vistas2noise.com↗

ChatTTS

ChatTTS is a conversational text-to-speech generative model designed to convert written dialogue into natural sounding audio. It functions as a multilingual speech synthesis framework capable of producing human-like audio across different languages and speaker profiles.

The system is distinguished by its ability to generate interactive dialogue with realistic vocal nuances. It utilizes a speech nuance controller to insert specific tokens that trigger non-verbal elements, such as laughter, pauses, and interjections, during the synthesis process.

The project includes a streaming audio generator that delivers speech incrementally to reduce latency. It further supports multi-speaker embeddings to maintain consistent vocal characteristics throughout a conversation.

Features

  • Conversational Audio Streams - Provides a generative model for natural, multi-speaker interactive dialogue and conversational audio streams.
  • Audio Tokenization - Converts raw audio waveforms into discrete numerical codes for processing by the language model.
  • Autoregressive Transformers - Implements an autoregressive transformer architecture to predict audio tokens for sequential speech generation.
  • Prosody Control Tokens - Inserts specialized control tokens to trigger non-verbal vocal behaviors like laughter and pauses.
  • Speaker Embeddings - Uses learned vector representations to maintain consistent vocal characteristics across different speakers.
  • Multilingual Synthesis - Provides a generative framework capable of producing human-like speech across multiple languages.
  • Text-to-Speech - Synthesizes natural human speech from written dialogue with human-like rhythms and nuances.
  • Latent Acoustic Mapping - Maps natural language input to a latent space to guide the generation of acoustic features.
  • Spoken Dialogue Generation - Creates spoken audio for multi-speaker interactions including realistic pauses and emotional cues.
  • Vocal - Refines audio output by adding human-like non-verbal elements such as laughter and interjections.
  • Speech Synthesis Markup Controls - Uses markup-like tokens to control prosody and insert fine-grained vocal elements like laughter.
  • Vocal Nuance Controllers - Implements a controller that inserts vocal tokens to trigger specific audio elements like laughter.
  • Chunked Audio Streaming - Streams audio output in chunks incrementally to minimize latency during speech generation.
  • Generative Audio Chunking - Yields audio waveform chunks sequentially during generation to enable immediate low-latency playback.
  • Additional AI Tools - Generative TTS model optimized for natural, expressive daily dialogue with fine-grained prosody control.
  • Core Models - The official repository for the core model implementation.
  • Speech Processing - Conversational text-to-speech for dialogue.
  • Speech Synthesis - Conversational text-to-speech for dialogue.

Historial de estrellas

Gráfico del historial de estrellas de 2noise/chatttsGráfico del historial de estrellas de 2noise/chattts

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a ChatTTS

Proyectos open-source similares, clasificados según cuántas características comparten con ChatTTS.
  • sparkaudio/spark-ttsAvatar de SparkAudio

    SparkAudio/Spark-TTS

    10,930Ver en GitHub↗

    Spark-TTS is a deep learning text-to-speech synthesis engine designed to convert written text into high-fidelity audio. It utilizes a transformer-based architecture and autoregressive sequence modeling to generate coherent speech, transforming linguistic input into natural-sounding waveforms through neural speech codec synthesis. The platform distinguishes itself through zero-shot voice cloning, which allows users to mimic a target speaker’s unique vocal identity using only a short reference audio sample without requiring additional model training. It also features cross-lingual phonetic mapp

    Python
    Ver en GitHub↗10,930
  • openbmb/voxcpmAvatar de OpenBMB

    OpenBMB/VoxCPM

    29,985Ver en GitHub↗

    VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice cloning tool and a synthetic voice designer, capable of generating natural speech across global languages and regional dialects using a GPU-accelerated audio generator. The project features a speech model fine-tuning framework that supports both full parameter updates and low-rank adaptation for customizing voice characteristics. It enables high-fidelity voice cloning from reference audio, including cross-lingual voice transfer and acoustic environment mimicry, as well as the crea

    Pythonaudiodeeplearningminicpm
    Ver en GitHub↗29,985
  • fishaudio/fish-speechAvatar de fishaudio

    fishaudio/fish-speech

    24,928Ver en GitHub↗

    This project is a generative speech synthesis engine that converts text into high-fidelity human speech. It utilizes a two-stage autoregressive transformer architecture that separates semantic token prediction from acoustic detail reconstruction to balance linguistic accuracy with audio quality. The system is designed to support multilingual output and conversational AI development, enabling the generation of context-aware speech that maintains flow across multiple dialogue turns. The platform distinguishes itself through a production-ready inference server that employs continuous batching to

    Pythonllamatransformertts
    Ver en GitHub↗24,928
  • boson-ai/higgs-audioAvatar de boson-ai

    boson-ai/higgs-audio

    7,919Ver en GitHub↗

    Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large language model architectures. It functions as a multilingual speech synthesizer capable of generating high-fidelity audio across different languages with control over emotional tone and prosody. The system includes a voice cloning tool that creates synthetic replicas of specific speakers from short audio samples without requiring extensive model training. It also provides a streaming audio API designed to deliver generated speech incrementally to minimize playback delay. The

    Python
    Ver en GitHub↗7,919
Ver las 30 alternativas a ChatTTS→

Preguntas frecuentes

¿Qué hace 2noise/chattts?

ChatTTS is a conversational text-to-speech generative model designed to convert written dialogue into natural sounding audio. It functions as a multilingual speech synthesis framework capable of producing human-like audio across different languages and speaker profiles.

¿Cuáles son las características principales de 2noise/chattts?

Las características principales de 2noise/chattts son: Conversational Audio Streams, Audio Tokenization, Autoregressive Transformers, Prosody Control Tokens, Speaker Embeddings, Multilingual Synthesis, Text-to-Speech, Latent Acoustic Mapping.

¿Qué alternativas de código abierto existen para 2noise/chattts?

Las alternativas de código abierto para 2noise/chattts incluyen: sparkaudio/spark-tts — Spark-TTS is a deep learning text-to-speech synthesis engine designed to convert written text into high-fidelity… openbmb/voxcpm — VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice… fishaudio/fish-speech — This project is a generative speech synthesis engine that converts text into high-fidelity human speech. It utilizes a… boson-ai/higgs-audio — Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large… microsoft/vibevoice — VibeVoice is a generative artificial intelligence platform designed for text-to-speech synthesis. It functions as a… index-tts/index-tts — Index-tts is a neural audio generation engine designed to convert written text into high-fidelity human speech. By…