awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

54 repositorios

Awesome GitHub RepositoriesSpeech Processing

Tools for speech recognition, synthesis, and audio analysis.

Explore 54 awesome GitHub repositories matching part of an awesome list · Speech Processing. Refine with filters or upvote what's useful.

Awesome Speech Processing GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • openai/whisperAvatar de openai

    openai/whisper

    102,828Ver en GitHub↗

    This project is a speech recognition and translation engine that utilizes a sequence-to-sequence transformer architecture to convert audio into text. It is built upon a weakly supervised learning framework, which leverages large-scale, unlabelled audio-transcript data to create generalized speech representations capable of performing simultaneous transcription, language identification, and translation. The system distinguishes itself through a unified multi-task modeling approach that shares token sequences across different objectives, allowing it to handle diverse languages and vocabularies

    Robust speech-to-text transcription model.

    Python
    Ver en GitHub↗102,828
  • rvc-boss/gpt-sovitsAvatar de RVC-Boss

    RVC-Boss/GPT-SoVITS

    58,724Ver en GitHub↗

    GPT-SoVITS is a text-to-speech synthesis engine and voice cloning toolkit designed for generating natural-sounding human speech. It functions as a neural audio processing pipeline that maps input text to high-fidelity audio waveforms, utilizing conditional variational autoencoders and flow-based decoders to ensure expressive output. The platform distinguishes itself through its ability to perform few-shot voice cloning and cross-lingual speech generation, allowing users to maintain a specific speaker's vocal identity and emotional delivery across multiple languages. By employing cross-modal l

    Few-shot voice conversion and TTS system.

    Pythontext-to-speechttsvits
    Ver en GitHub↗58,724
  • microsoft/vibevoiceAvatar de microsoft

    microsoft/VibeVoice

    49,394Ver en GitHub↗

    VibeVoice is a generative artificial intelligence platform designed for text-to-speech synthesis. It functions as a neural audio generation framework that converts written text into natural-sounding spoken audio, specifically engineered to maintain consistent vocal characteristics and narrative prosody across extended passages of content. The system distinguishes itself through its ability to generate long-form conversational speech while preserving speaker identity and linguistic content. By utilizing latent space disentanglement, the model separates speaker traits from the input text, allow

    Voice interaction and synthesis framework.

    Python
    Ver en GitHub↗49,394
  • 2noise/chatttsAvatar de 2noise

    2noise/ChatTTS

    39,464Ver en GitHub↗

    ChatTTS is a conversational text-to-speech generative model designed to convert written dialogue into natural sounding audio. It functions as a multilingual speech synthesis framework capable of producing human-like audio across different languages and speaker profiles. The system is distinguished by its ability to generate interactive dialogue with realistic vocal nuances. It utilizes a speech nuance controller to insert specific tokens that trigger non-verbal elements, such as laughter, pauses, and interjections, during the synthesis process. The project includes a streaming audio generato

    Conversational text-to-speech for dialogue.

    Pythonagentchatchatgpt
    Ver en GitHub↗39,464
  • suno-ai/barkAvatar de suno-ai

    suno-ai/bark

    39,159Ver en GitHub↗

    Bark is a generative audio engine and machine learning inference library designed to convert written text into high-fidelity speech and sound effects. It functions as a text-to-audio transformer, utilizing multi-stage neural network architectures to map semantic input tokens into detailed audio codebooks for synthesis. The system distinguishes itself through a hierarchical transformer stacking approach that separates semantic understanding from acoustic realization. By employing autoregressive token prediction and vector quantized codebook mapping, the engine bridges linguistic and sonic doma

    Transformer-based text-to-audio model.

    Jupyter Notebook
    Ver en GitHub↗39,159
  • myshell-ai/openvoiceAvatar de myshell-ai

    myshell-ai/OpenVoice

    36,720Ver en GitHub↗

    OpenVoice is a multilingual text-to-speech framework and voice cloning AI model designed for high-fidelity voice replication and low-latency audio generation. It functions as an instant speech synthesis engine that converts text to audio while replicating a specific speaker's tone and color. The system is distinguished by its ability to perform cross-lingual cloning, allowing the vocal characteristics of a reference speaker to be applied to speech in different languages regardless of the original training data. It utilizes a decoupled representation to separate the physical identity of a voic

    Instant voice cloning and speech synthesis.

    Pythontext-to-speechttsvoice-clone
    Ver en GitHub↗36,720
  • openbmb/voxcpmAvatar de OpenBMB

    OpenBMB/VoxCPM

    29,985Ver en GitHub↗

    VoxCPM is a multilingual speech synthesis system and text-to-speech inference server. It functions as an AI voice cloning tool and a synthetic voice designer, capable of generating natural speech across global languages and regional dialects using a GPU-accelerated audio generator. The project features a speech model fine-tuning framework that supports both full parameter updates and low-rank adaptation for customizing voice characteristics. It enables high-fidelity voice cloning from reference audio, including cross-lingual voice transfer and acoustic environment mimicry, as well as the crea

    Multimodal speech and language model.

    Pythonaudiodeeplearningminicpm
    Ver en GitHub↗29,985
  • fishaudio/fish-speechAvatar de fishaudio

    fishaudio/fish-speech

    24,928Ver en GitHub↗

    This project is a generative speech synthesis engine that converts text into high-fidelity human speech. It utilizes a two-stage autoregressive transformer architecture that separates semantic token prediction from acoustic detail reconstruction to balance linguistic accuracy with audio quality. The system is designed to support multilingual output and conversational AI development, enabling the generation of context-aware speech that maintains flow across multiple dialogue turns. The platform distinguishes itself through a production-ready inference server that employs continuous batching to

    Advanced speech synthesis and generation model.

    Pythonllamatransformertts
    Ver en GitHub↗24,928
  • funaudiollm/cosyvoiceAvatar de FunAudioLLM

    FunAudioLLM/CosyVoice

    21,673Ver en GitHub↗

    CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus

    Multilingual speech generation and cloning model.

    Pythonaudio-generationcantonesechatbot
    Ver en GitHub↗21,673
  • nari-labs/diaAvatar de nari-labs

    nari-labs/dia

    19,324Ver en GitHub↗

    Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt

    Voice interaction and speech synthesis framework.

    Pythonaiopen-weighttext-to-speech
    Ver en GitHub↗19,324
  • index-tts/index-ttsAvatar de index-tts

    index-tts/index-tts

    18,851Ver en GitHub↗

    Index-tts is a neural audio generation engine designed to convert written text into high-fidelity human speech. By utilizing deep learning models and phoneme-based sequence modeling, the system transforms text into natural-sounding audio waveforms suitable for a variety of accessibility and media applications. The platform functions as a server-side inference pipeline that provides a programmatic interface for integrating voice generation into external applications. It distinguishes itself through asynchronous audio streaming, which buffers and delivers generated speech chunks in real time to

    Text-to-speech synthesis framework.

    Pythonbigvgancross-lingualindextts
    Ver en GitHub↗18,851
  • swivid/f5-ttsAvatar de SWivid

    SWivid/F5-TTS

    14,798Ver en GitHub↗

    F5-TTS is a text-to-speech system that utilizes a flow matching engine and diffusion transformers to generate fluent synthetic speech. It functions as a multilingual speech synthesizer and neural training framework, providing tools for voice cloning and high-performance inference serving. The project distinguishes itself through a voice cloning toolkit capable of mimicking specific speaker characteristics and tones from reference audio clips. It supports cross-lingual generation, allowing for the synthesis of audio across various global languages or the mixing of multiple languages within a s

    Fast and high-quality text-to-speech synthesis.

    Python
    Ver en GitHub↗14,798
  • sparkaudio/spark-ttsAvatar de SparkAudio

    SparkAudio/Spark-TTS

    10,930Ver en GitHub↗

    Spark-TTS is a deep learning text-to-speech synthesis engine designed to convert written text into high-fidelity audio. It utilizes a transformer-based architecture and autoregressive sequence modeling to generate coherent speech, transforming linguistic input into natural-sounding waveforms through neural speech codec synthesis. The platform distinguishes itself through zero-shot voice cloning, which allows users to mimic a target speaker’s unique vocal identity using only a short reference audio sample without requiring additional model training. It also features cross-lingual phonetic mapp

    Efficient text-to-speech synthesis model.

    Python
    Ver en GitHub↗10,930
  • kittenml/kittenttsAvatar de KittenML

    KittenML/KittenTTS

    10,044Ver en GitHub↗

    KittenTTS is a neural text-to-speech engine and text-to-audio synthesis tool that converts written text into spoken audio using lightweight neural network models. It functions as both a speech synthesizer and an audio file generator, producing spoken audio for offline playback. The system includes a text normalization processor that expands numbers and abbreviations into full spoken words to improve the naturalness of the synthesized speech. It supports diverse voice options and provides the ability to adjust playback speed.

    Lightweight text-to-speech synthesis.

    Python
    Ver en GitHub↗10,044
  • rany2/edge-ttsAvatar de rany2

    rany2/edge-tts

    10,041Ver en GitHub↗

    edge-tts is a command line interface and text-to-speech engine that converts written text into audio files using the Microsoft Edge online synthesis service. It functions as a client for generating high-quality speech and managing the conversion of text to audio. The project provides utilities for generating synchronized SRT subtitle files by tracking word and sentence boundaries during synthesis. It also includes a voice profile discovery system to browse a catalog of available synthetic voices based on gender and personality traits. Users can customize vocal characteristics by adjusting th

    Library for using Microsoft Edge's TTS service.

    Pythonspeech-synthesistext-to-speechtts
    Ver en GitHub↗10,041
  • boson-ai/higgs-audioAvatar de boson-ai

    boson-ai/higgs-audio

    7,919Ver en GitHub↗

    Higgs-audio is a generative text-to-speech engine that transforms text into natural conversational speech using large language model architectures. It functions as a multilingual speech synthesizer capable of generating high-fidelity audio across different languages with control over emotional tone and prosody. The system includes a voice cloning tool that creates synthetic replicas of specific speakers from short audio samples without requiring extensive model training. It also provides a streaming audio API designed to deliver generated speech incrementally to minimize playback delay. The

    Audio generation and synthesis framework.

    Python
    Ver en GitHub↗7,919
  • k2-fsa/omnivoiceAvatar de k2-fsa

    k2-fsa/OmniVoice

    7,689Ver en GitHub↗

    OmniVoice is a Python-based open-source project hosted under the k2-fsa organization. The repository currently does not contain any documented features or structured documentation, making it difficult to determine its specific purpose or capabilities at this time.

    Unified voice interaction and synthesis.

    Python
    Ver en GitHub↗7,689
  • bytedance/megatts3Avatar de bytedance

    bytedance/MegaTTS3

    6,066Ver en GitHub↗

    MegaTTS3 is a bilingual speech synthesis system that generates natural-sounding speech in Chinese and English, including seamless code-switching within a single utterance. It functions as a text-to-speech engine, voice cloning system, and speech-to-text alignment tool, built around an acoustic latent compression model that encodes high-resolution audio into compact representations for efficient processing. The system distinguishes itself through accent intensity control, allowing adjustment of a speaker's accent strength in generated speech, and voice cloning from short audio samples for pers

    High-performance text-to-speech model.

    Pythonresearch
    Ver en GitHub↗6,066
  • hexgrad/kokoroAvatar de hexgrad

    hexgrad/kokoro

    5,729Ver en GitHub↗

    Kokoro is a lightweight neural text-to-speech engine that converts written text into spoken audio using a compact model designed for fast inference. It supports multiple languages through language-specific grapheme-to-phoneme conversion pipelines, and offers voice profile selection to change the character of the generated speech. The engine provides GPU acceleration on Apple Silicon hardware by setting a single environment variable, enabling faster inference on Mac M-series machines. It also includes pattern-based text segmentation, allowing input text to be split at user-defined delimiters t

    Lightweight and fast text-to-speech model.

    JavaScript
    Ver en GitHub↗5,729
  • kyutai-labs/delayed-streams-modelingAvatar de kyutai-labs

    kyutai-labs/delayed-streams-modeling

    2,955Ver en GitHub↗

    Kyutai's Speech-To-Text and Text-To-Speech models based on the Delayed Streams Modeling framework.

    Real-time speech processing and modeling.

    Python
    Ver en GitHub↗2,955
Ant.123Siguiente
  1. Home
  2. Part of an Awesome List
  3. Media & Communication
  4. Speech Processing