awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
KittenML avatar

KittenML/KittenTTS

0
View on GitHub↗
10,044 stars·525 forks·Python·apache-2.0·13 views

KittenTTS

KittenTTS is a neural text-to-speech engine and text-to-audio synthesis tool that converts written text into spoken audio using lightweight neural network models. It functions as both a speech synthesizer and an audio file generator, producing spoken audio for offline playback.

The system includes a text normalization processor that expands numbers and abbreviations into full spoken words to improve the naturalness of the synthesized speech. It supports diverse voice options and provides the ability to adjust playback speed.

Features

  • Neural Text-to-Speech Engines - Uses lightweight neural network models to map text directly to audio waveforms for natural speech synthesis.
  • Text-to-Audio Synthesis - Converts written text into audio samples using lightweight models with adjustable playback speed.
  • Text-to-Speech - Synthesizes spoken audio from written text using a neural network model optimized for speed.
  • Speech Synthesis Formatters - Prepares raw text by expanding abbreviations and numbers to ensure high-quality synthesized speech.
  • Text-to-Speech Normalizers - Processes raw text to expand numbers and abbreviations into full spoken words for more natural synthesis.
  • Automated Content Creation Tools - Generates audio files from text scripts to create voiceovers without manual recording.
  • Lightweight Voice Generation - Produces spoken audio with adjustable speed using models optimized for low-resource hardware.
  • Model Quantization - Uses model quantization to reduce precision, lowering memory usage and increasing inference speed on consumer hardware.
  • Prosody Controls - Allows adjustment of speech speed and tone by modifying input variables during the inference process.
  • Audio Generation - Functions as a generator that writes synthesized speech directly to audio files.
  • Audio Exporters - Provides the ability to serialize synthesized audio buffers into standard files for permanent storage.
  • Generative Audio Chunking - Implements sequential yielding of audio chunks during synthesis to enable low-latency playback.
  • Audio and Video Processing - Provides lightweight, CPU-friendly text-to-speech synthesis.
  • Audio Video Processing - Listed in the “Audio Video Processing” section of the Awesome Python awesome list.
  • Speech Processing - Lightweight text-to-speech synthesis.
  • Speech Synthesis - Lightweight speech synthesis for various applications.

Star history

Star history chart for kittenml/kittenttsStar history chart for kittenml/kittentts

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Open-source alternatives to KittenTTS

Similar open-source projects, ranked by how many features they share with KittenTTS.
  • nari-labs/dianari-labs avatar

    nari-labs/dia

    19,324View on GitHub↗

    Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt

    Pythonaiopen-weighttext-to-speech
    View on GitHub↗19,324
  • index-tts/index-ttsindex-tts avatar

    index-tts/index-tts

    18,851View on GitHub↗

    Index-tts is a neural audio generation engine designed to convert written text into high-fidelity human speech. By utilizing deep learning models and phoneme-based sequence modeling, the system transforms text into natural-sounding audio waveforms suitable for a variety of accessibility and media applications. The platform functions as a server-side inference pipeline that provides a programmatic interface for integrating voice generation into external applications. It distinguishes itself through asynchronous audio streaming, which buffers and delivers generated speech chunks in real time to

    Pythonbigvgancross-lingualindextts
    View on GitHub↗18,851
  • hexgrad/kokorohexgrad avatar

    hexgrad/kokoro

    5,729View on GitHub↗

    Kokoro is a lightweight neural text-to-speech engine that converts written text into spoken audio using a compact model designed for fast inference. It supports multiple languages through language-specific grapheme-to-phoneme conversion pipelines, and offers voice profile selection to change the character of the generated speech. The engine provides GPU acceleration on Apple Silicon hardware by setting a single environment variable, enabling faster inference on Mac M-series machines. It also includes pattern-based text segmentation, allowing input text to be split at user-defined delimiters t

    JavaScript
    View on GitHub↗5,729
  • boson-ai/higgs-audioboson-ai avatar

    boson-ai/higgs-audio

    7,919View on GitHub↗

    Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large language model architectures. It functions as a multilingual speech synthesizer capable of generating high-fidelity audio across different languages with control over emotional tone and prosody. The system includes a voice cloning tool that creates synthetic replicas of specific speakers from short audio samples without requiring extensive model training. It also provides a streaming audio API designed to deliver generated speech incrementally to minimize playback delay. The

    Python
    View on GitHub↗7,919
See all 30 alternatives to KittenTTS→

Frequently asked questions

What does kittenml/kittentts do?

KittenTTS is a neural text-to-speech engine and text-to-audio synthesis tool that converts written text into spoken audio using lightweight neural network models. It functions as both a speech synthesizer and an audio file generator, producing spoken audio for offline playback.

What are the main features of kittenml/kittentts?

The main features of kittenml/kittentts are: Neural Text-to-Speech Engines, Text-to-Audio Synthesis, Text-to-Speech, Speech Synthesis Formatters, Text-to-Speech Normalizers, Automated Content Creation Tools, Lightweight Voice Generation, Model Quantization.

What are some open-source alternatives to kittenml/kittentts?

Open-source alternatives to kittenml/kittentts include: nari-labs/dia — Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of… index-tts/index-tts — Index-tts is a neural audio generation engine designed to convert written text into high-fidelity human speech. By… hexgrad/kokoro — Kokoro is a lightweight neural text-to-speech engine that converts written text into spoken audio using a compact… boson-ai/higgs-audio — Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large… openbmb/voxcpm — VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice… myshell-ai/openvoice — OpenVoice is a multilingual text-to-speech framework and voice cloning AI model designed for high-fidelity voice…