awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

Modelos de síntesis y reconocimiento de voz

Clasificación actualizada el 30 jun 2026

For una herramienta de código abierto para síntesis y reconocimiento de voz, the strongest matches are paddlepaddle/paddlespeech (PaddleSpeech is a comprehensive speech toolkit providing pre-trained ASR), nvidia/nemo (NeMo is a comprehensive open-source AI framework that includes) and espnet/espnet (ESPnet is a comprehensive open-source speech processing toolkit that). plachtaa/vall-e-x and blaizzy/mlx-audio round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.

Librerías open-source y modelos preentrenados para convertir audio hablado a texto y generar voz humana sintética.

Modelos de síntesis y reconocimiento de voz

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • paddlepaddle/paddlespeechAvatar de PaddlePaddle

    PaddlePaddle/PaddleSpeech

    12,626Ver en GitHub↗

    PaddleSpeech is a comprehensive toolkit of neural models for speech recognition, synthesis, and translation built on the PaddlePaddle deep learning framework. It provides a collection of frameworks and tools for converting spoken audio into written text, synthesizing natural audio from text, and performing direct speech translation. The toolkit includes specialized capabilities for keyword spotting to detect trigger words and speaker verification systems that extract unique voiceprints to identify and distinguish between individuals. It also features end-to-end translation tools that map audi

    PaddleSpeech is a comprehensive speech toolkit providing pre-trained ASR (including Whisper and Wav2Vec2) and TTS models with multilingual support, voice cloning, and real-time streaming, making it an excellent fit for self-hosted integration.

    PythonAutomatic Speech RecognitionNeural VocodersSelf-Supervised Speech Representations
    Ver en GitHub↗12,626
  • nvidia/nemoAvatar de NVIDIA

    NVIDIA/NeMo

    17,394Ver en GitHub↗

    NeMo is a multimodal AI framework and toolkit designed for the development, training, and scaling of large language models, generative AI systems, and speech-based models. It functions as an automatic speech recognition toolkit, a text-to-speech engine, and a framework for building models that process and generate combinations of text, image, and audio data. The project serves as a conversational AI orchestrator capable of managing real-time, interruptible voice interactions. It provides specialized workflows for speech translation, converting spoken audio from one language into text or speec

    NeMo is a comprehensive open-source AI framework that includes pre-trained ASR and TTS models, supports multilingual speech, real-time inference, fine-tuning, and voice cloning, making it a strong fit for integrating speech recognition and synthesis into applications.

    PythonAutomatic Speech RecognitionAutomatic Speech Recognition
    Ver en GitHub↗17,394
  • espnet/espnetAvatar de espnet

    espnet/espnet

    9,861Ver en GitHub↗

    ESPnet is a comprehensive speech processing toolkit and PyTorch-based trainer designed for building end-to-end speech recognition, synthesis, and translation models. It provides a structured framework for developing automatic speech recognition systems using transducer and encoder-decoder architectures, alongside engines for text-to-speech synthesis and speech translation pipelines. The project distinguishes itself through a recipe-based workflow execution system that ensures experimental reproducibility by running standardized sequences of scripts for data preparation and model training. It

    ESPnet is a comprehensive open-source speech processing toolkit that provides both pre-trained ASR and TTS models (with multilingual support, fine-tuning, streaming inference, and voice conversion), making it a ready-to-use self-hostable solution for building and deploying speech recognition and synthesis.

    PythonAutomatic Speech RecognitionSpeech Model Fine-TuningSpeech Recognition Models
    Ver en GitHub↗9,861
  • plachtaa/vall-e-xAvatar de Plachtaa

    Plachtaa/VALL-E-X

    7,939Ver en GitHub↗

    VALL-E-X is a neural speech synthesis framework and zero-shot text-to-speech engine. It functions as a multilingual synthesizer capable of generating natural human speech with control over emotion, pitch, and prosody. The project specializes in zero-shot voice cloning and cross-lingual voice replication, allowing the system to produce personalized speech in multiple target languages using short audio samples without additional training. It further enables cross-language accent manipulation and the ability to match the emotional tone and acoustic environment of a provided prompt. The implemen

    VALL-E-X is a zero-shot multilingual text-to-speech engine with voice cloning and emotional control, matching your TTS needs, but it does not include automatic speech recognition (ASR) models, so it covers only one half of the combined ASR and TTS requirement.

    PythonMultilingual Speech ModelsVoice Cloning
    Ver en GitHub↗7,939
  • blaizzy/mlx-audioAvatar de Blaizzy

    Blaizzy/mlx-audio

    5,994Ver en GitHub↗

    mlx-audio is an audio processing toolkit built on Apple MLX that provides speech transcription, text-to-speech synthesis, voice cloning, and audio source separation using local models. It offers an OpenAI-compatible REST API and web interface for running audio generation and transcription tasks, enabling drop-in integration with existing tools that follow that endpoint structure. The toolkit supports text-prompted audio source separation, allowing specific sounds to be isolated from mixed recordings based on natural language descriptions. It also provides voice cloning from a short reference

    mlx-audio is an audio-processing toolkit that provides both speech transcription (ASR) and text-to-speech (TTS) via local models, plus voice cloning and an OpenAI-compatible API for self-hosting — fitting your need for open-source ASR/TTS libraries, though it is currently limited to Apple Silicon hardware.

    PythonVoice Cloning
    Ver en GitHub↗5,994
  • openai/whisperAvatar de openai

    openai/whisper

    102,828Ver en GitHub↗

    This project is a speech recognition and translation engine that utilizes a sequence-to-sequence transformer architecture to convert audio into text. It is built upon a weakly supervised learning framework, which leverages large-scale, unlabelled audio-transcript data to create generalized speech representations capable of performing simultaneous transcription, language identification, and translation. The system distinguishes itself through a unified multi-task modeling approach that shares token sequences across different objectives, allowing it to handle diverse languages and vocabularies

    Whisper is a state-of-the-art open-source speech recognition model supporting multilingual transcription and translation, which directly addresses the ASR side of the query but does not include any text-to-speech synthesis capabilities, so it partially meets the combined intent.

    PythonAutomatic Speech Recognition
    Ver en GitHub↗102,828
  • microsoft/vibevoiceAvatar de microsoft

    microsoft/VibeVoice

    49,394Ver en GitHub↗

    VibeVoice is a generative artificial intelligence platform designed for text-to-speech synthesis. It functions as a neural audio generation framework that converts written text into natural-sounding spoken audio, specifically engineered to maintain consistent vocal characteristics and narrative prosody across extended passages of content. The system distinguishes itself through its ability to generate long-form conversational speech while preserving speaker identity and linguistic content. By utilizing latent space disentanglement, the model separates speaker traits from the input text, allow

    VibeVoice is a generative TTS model built for long-form speech synthesis with voice cloning and speaker consistency, which directly addresses the text-to-speech side of your search, but it does not include any automatic speech recognition capabilities.

    PythonNeural Vocoders
    Ver en GitHub↗49,394
  • k2-fsa/sherpa-onnxAvatar de k2-fsa

    k2-fsa/sherpa-onnx

    13,017Ver en GitHub↗

    Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device voice synthesis tool, and a speaker identification framework. It is designed as a cross-platform speech API that enables speech-to-text, text-to-speech, and speaker verification tasks to be executed locally on a device without requiring network access. The project is distinguished by its ability to perform zero-shot voice cloning and speaker diarization on-device. It supports a wide range of hardware accelerations, including GPU and various NPU architectures, and provides a Web

    Sherpa-ONNX is a local speech toolkit with on-device ASR, TTS, and zero-shot voice cloning, making it a solid fit for self-hosting and integration, though it does not explicitly advertise fine-tuning or broad multilingual model coverage.

    C++Voice Cloning
    Ver en GitHub↗13,017
  • 2noise/chatttsAvatar de 2noise

    2noise/ChatTTS

    39,464Ver en GitHub↗

    ChatTTS is a conversational text-to-speech generative model designed to convert written dialogue into natural sounding audio. It functions as a multilingual speech synthesis framework capable of producing human-like audio across different languages and speaker profiles. The system is distinguished by its ability to generate interactive dialogue with realistic vocal nuances. It utilizes a speech nuance controller to insert specific tokens that trigger non-verbal elements, such as laughter, pauses, and interjections, during the synthesis process. The project includes a streaming audio generato

    ChatTTS is a multilingual text-to-speech generative model with nuanced vocal control and streaming audio, fitting the TTS side of this search, but it does not provide automatic speech recognition (ASR) so it covers only half of the requested capabilities.

    PythonConversational Audio StreamsAudio TokenizationAutoregressive Transformers
    Ver en GitHub↗39,464
  • huggingface/transformersAvatar de huggingface

    huggingface/transformers

    161,630Ver en GitHub↗

    Transformers is a comprehensive library for machine learning that provides a unified interface for training, fine-tuning, and deploying transformer-based models. It supports a wide range of tasks, including text classification, language modeling, question answering, and sequence-to-sequence translation, while offering specialized architectures for both text and vision processing. The framework includes tools for managing the entire model lifecycle, from data preprocessing and tokenization to distributed training and inference. The library features extensive support for model optimization and

    Hugging Face Transformers is a comprehensive library that provides a unified interface for thousands of pre-trained models, including widely-used ASR models (e.g., Whisper, Wav2Vec2) and TTS models (e.g., VITS, Tacotron2), with support for fine-tuning, multilingual models, and self-hosted deployment, making it a strong fit for both speech recognition and synthesis needs.

    PythonAPI FrameworksByte Pair EncodingsHybrid
    Ver en GitHub↗161,630
  • argmaxinc/whisperkitAvatar de argmaxinc

    argmaxinc/WhisperKit

    5,639Ver en GitHub↗

    WhisperKit is an Apple-platform library for running OpenAI's Whisper speech recognition with added synthesis capabilities, fitting the request for self-hostable ASR/TTS integration, though it is limited to Apple devices and the TTS support is less documented.

    SwiftOn-Device Speech-to-Text SDKsSpeech to Text TranscriptionAudio Streaming Pipelines
    Ver en GitHub↗5,639
Compara los 10 mejores de un vistazo
RepositorioEstrellasLenguajeLicenciaÚltimo push
paddlepaddle/paddlespeech12.6KPythonApache-2.021 jun 2026
nvidia/nemo17.4KPythonApache-2.017 jun 2026
espnet/espnet9.9KPythonApache-2.017 jun 2026
plachtaa/vall-e-x7.9KPythonMIT11 feb 2024
blaizzy/mlx-audio6KPythonmit19 feb 2026
openai/whisper102.8KPythonMIT15 abr 2026
microsoft/vibevoice49.4KPythonMIT6 may 2026
k2-fsa/sherpa-onnx13KC++Apache-2.015 jun 2026
2noise/chattts39.5KPythonAGPL-3.010 abr 2026
huggingface/transformers161.6KPythonApache-2.016 jun 2026

Related searches

  • un motor de texto a voz autohospedado
  • un motor para reconocimiento de voz offline
  • toolkit para clonación de voz con IA
  • una herramienta open source para clonación de voz
  • an open source real time voice changer
  • Streaming speech recognition
  • una librería para reconocimiento de voz en streaming
  • una aplicación para transcribir notas de voz