awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoServidor MCPAcerca deCómo clasificamosPrensa
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

13 repositorios

Awesome GitHub RepositoriesAudio Dataset Preprocessing

Tools for cleaning and standardizing raw audio and open-source speech datasets for ML training.

Distinct from Dataset Preprocessing Tools: Focuses on audio-specific cleaning and unification, whereas the parent is a general ML preprocessing utility.

Explore 13 awesome GitHub repositories matching artificial intelligence & ml · Audio Dataset Preprocessing. Refine with filters or upvote what's useful.

Awesome Audio Dataset Preprocessing GitHub Repositories

Encuentra los mejores repositorios con IA.Buscaremos los repositorios que mejor coincidan usando IA.
  • espnet/espnetAvatar de espnet

    espnet/espnet

    9,861Ver en GitHub↗

    ESPnet is a comprehensive speech processing toolkit and PyTorch-based trainer designed for building end-to-end speech recognition, synthesis, and translation models. It provides a structured framework for developing automatic speech recognition systems using transducer and encoder-decoder architectures, alongside engines for text-to-speech synthesis and speech translation pipelines. The project distinguishes itself through a recipe-based workflow execution system that ensures experimental reproducibility by running standardized sequences of scripts for data preparation and model training. It

    Converts raw audio files into structured manifests required for model training and evaluation.

    Python
    Ver en GitHub↗9,861
  • open-mmlab/amphionAvatar de open-mmlab

    open-mmlab/Amphion

    9,844Ver en GitHub↗

    Amphion is an audio generation toolkit designed for the research and development of models that synthesize speech, music, and environmental sound effects. It provides a standardized framework for reproducible audio synthesis, incorporating a text-to-speech engine and a voice conversion framework. The project specializes in transforming audio identities, allowing for the modification of speaker accents and voice identities while preserving original rhythm and style. It also includes capabilities for singing voice synthesis and the generation of environmental soundscapes from text descriptions

    Unifies the cleaning and preparation of various open-source audio datasets and raw speech data.

    Pythonaudio-generationaudio-synthesisaudioldm
    Ver en GitHub↗9,844
  • voicepaw/so-vits-svc-forkAvatar de voicepaw

    voicepaw/so-vits-svc-fork

    9,318Ver en GitHub↗

    This project is an AI singing voice conversion system and vocal processor used for training generative voice models and converting vocal recordings or live input into a target voice. It functions as a VITS model trainer and a real-time voice changer that transforms vocal timbre and pitch to change the identity of a singer. The system provides a graphical management dashboard for controlling training hyperparameters and voice conversion presets. It supports low-latency audio streaming for live microphone input and employs pitch estimation to ensure precise matching between source and target vo

    Provides tools for cleaning, segmenting, and standardizing raw audio recordings for ML training.

    Pythoncontentvecdeep-learninggan
    Ver en GitHub↗9,318
  • dusty-nv/jetson-inferenceAvatar de dusty-nv

    dusty-nv/jetson-inference

    8,734Ver en GitHub↗

    jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput. The project distinguishes itself through high-performance streaming analytics and the ability to execute concurrent AI pipelines on auto-grade silicon. It provides specialized support for multi-sensor stream processing, utilizing zero-copy data transport to load camera frames directly into GPU memory. The codebase covers a broad surface of capabiliti

    Transcribes and filters speech data using automatic speech recognition to prepare high-quality audio datasets.

    C++caffecomputer-visiondeep-learning
    Ver en GitHub↗8,734
  • nl8590687/asrt_speechrecognitionAvatar de nl8590687

    nl8590687/ASRT_SpeechRecognition

    8,375Ver en GitHub↗

    This project is a Chinese automatic speech recognition framework and deep learning system designed to convert spoken Chinese audio into written text. It functions as a toolkit for training, evaluating, and deploying speech-to-text models, utilizing a specialized pinyin-to-text converter that transforms phonetic sequences into Chinese characters using a probability graph model. The system is distinguished by its deployment flexibility, offering a dockerized recognition server that provides transcription capabilities as a remote API. It supports high-performance streaming through a gRPC speech-

    Implements tools for cleaning and standardizing raw audio datasets specifically for machine learning training.

    Pythonasrtchinese-speech-recognitioncnn
    Ver en GitHub↗8,375
  • snakers4/silero-vadAvatar de snakers4

    snakers4/silero-vad

    8,209Ver en GitHub↗

    Silero VAD is a voice activity detection model and deep learning speech classifier designed to distinguish human speech from silence across diverse languages and noisy environments. It functions as a pre-trained neural network capable of identifying speech segments within both static audio recordings and real-time data streams. The project includes a language identification tool for classifying spoken languages and a framework for fine-tuning audio models. It provides utilities for optimizing detection thresholds using validation datasets and retraining the model with custom labeled audio to

    Isolates and merges speech segments from a recording to remove silence before transcription.

    Pythononnxonnx-runtimeonnxruntime
    Ver en GitHub↗8,209
  • facebookresearch/auglyAvatar de facebookresearch

    facebookresearch/AugLy

    5,086Ver en GitHub↗

    AugLy es una biblioteca de aumento de datos multimodal y aumentador de conjuntos de datos de machine learning. Proporciona un sistema para generar variaciones sintéticas de datos de entrenamiento a través de conjuntos de datos de audio, imagen, texto y video para aumentar la diversidad de muestras y mejorar la robustez del modelo. La biblioteca funciona como un simulador de ruido multimedia, diseñado específicamente para imitar capturas de usuarios del mundo real superponiendo plantillas de redes sociales y artefactos de internet sobre los medios. Incluye un rastreador de procedencia de datos para registrar las transformaciones específicas y los niveles de intensidad aplicados a cada pieza de datos aumentados. La herramienta cubre una amplia gama de capacidades de expansión de conjuntos de datos, incluyendo transformaciones lingüísticas para texto, transformaciones temporales y visuales para video, y transformaciones sónicas para audio.

    Applies transformations to audio files to create more varied training samples for sound recognition or processing models.

    Python
    Ver en GitHub↗5,086
  • microsoft/muzicAvatar de microsoft

    microsoft/muzic

    4,928Ver en GitHub↗

    Muzic es una plataforma y framework de deep learning para el análisis, composición y síntesis de música impulsada por IA. Funciona como un framework de generación de música y herramienta de análisis, utilizando modelos de lenguaje grandes y agentes autónomos para orquestar la creación e interpretación de música simbólica y de audio. El proyecto se distingue por sus capacidades multimodales, mapeando el lenguaje natural y la música simbólica en un espacio de incrustación (embedding) conjunto compartido para clasificación zero-shot y recuperación de información. Emplea una variedad de arquitecturas especializadas, incluyendo frameworks de difusión para síntesis de audio, mecanismos de atención de grano dual para consistencia estructural de secuencias largas y un sistema híbrido que combina reglas de teoría musical con redes neuronales. La plataforma cubre una amplia gama de capacidades, incluyendo la generación de secuencias MIDI a partir de texto y letras, síntesis de voz cantada neuronal y transcripción automatizada de letras. También proporciona herramientas para el modelado de estructuras musicales, generación simbólica basada en atributos y la orquestación de herramientas musicales externas a través de agentes autónomos. Las utilidades de soporte incluyen pipelines de ingeniería de datos para la binarización de MIDI a gran escala, codificación de conjuntos de datos y procesamiento de señales de audio para la extracción de notas de melodía y alineación de voz a fonema.

    Provides tools for cleaning and converting raw MIDI and audio files into formats suitable for ML training.

    Pythonai-musicdeep-learningmusic
    Ver en GitHub↗4,928
  • buriburisuri/speech-to-text-wavenetAvatar de buriburisuri

    buriburisuri/speech-to-text-wavenet

    4,007Ver en GitHub↗

    Este proyecto es un framework de aprendizaje profundo diseñado para la transcripción de voz a texto de extremo a extremo. Utiliza la arquitectura de red neuronal WaveNet para procesar la entrada de audio hablado y generar transcripciones de texto escrito, aprovechando la clasificación temporal conexionista (CTC) para mapear secuencias de audio de longitud variable a salidas a nivel de carácter. El sistema se distingue por una tubería de entrenamiento integral que soporta la ejecución distribuida a través de múltiples unidades de procesamiento gráfico (GPU). Incluye utilidades especializadas para la aumentación de datos de audio y la transformación de archivos de audio crudos en formatos binarios optimizados, lo que minimiza la latencia de entrada y salida de disco durante el entrenamiento de modelos a gran escala. El software proporciona un entorno completo para gestionar flujos de trabajo de aprendizaje automático, incluyendo herramientas para calcular métricas de pérdida para monitorear la convergencia y precisión del modelo. Todos los componentes, incluyendo el motor de reconocimiento y las tuberías de entrenamiento, están diseñados para el despliegue dentro de entornos contenedorizados para asegurar una ejecución consistente en diversos sistemas host.

    Transforms raw audio files into optimized feature formats to accelerate machine learning training and reduce disk input bottlenecks.

    Python
    Ver en GitHub↗4,007
  • tensorspeech/tensorflowttsAvatar de TensorSpeech

    TensorSpeech/TensorflowTTS

    3,993Ver en GitHub↗

    TensorFlowTTS is a neural speech synthesis framework used to convert text into high-fidelity audio waveforms. It provides a toolkit for training and fine-tuning sequence-to-sequence or generative adversarial network architectures to produce natural sounding speech. The system includes neural vocoder implementations that transform intermediate acoustic representations into final audio waveforms. It also features playback speed control to adjust the rate of synthesized speech output. The framework covers the end-to-end pipeline for speech synthesis, including audio data preprocessing to create

    Provides utilities to convert raw audio and transcriptions into normalized mel spectrograms for ML training.

    Python
    Ver en GitHub↗3,993
  • stability-ai/stable-audio-toolsAvatar de Stability-AI

    Stability-AI/stable-audio-tools

    3,790Ver en GitHub↗

    Stable-audio-tools is a toolkit for training and deploying latent diffusion models for high-fidelity audio synthesis. It provides a framework for generating audio by iteratively refining noise within a compressed latent space, using specialized encoders to preserve temporal and spectral features of the audio signal. The project features a system for adapting pre-trained audio checkpoints to new datasets through modular initialization and configuration files. It includes utilities for weight extraction and inference model export, which remove training metadata and optimizer states to create li

    Integrates audio data from local directories or cloud stores for use in machine learning training pipelines.

    Python
    Ver en GitHub↗3,790
  • carykh/jumpcutterAvatar de carykh

    carykh/jumpcutter

    3,141Ver en GitHub↗

    Jumpcutter is an audio-based video cutter and automatic editor designed to eliminate dead air from video files. It functions as a utility that condenses footage by detecting and removing silent sections based on audio track analysis. The tool utilizes FFmpeg to automatically identify quiet gaps and strip them from recordings. This process focuses on removing silent video sections to create faster-paced content without the need for manual editing. The system operates by calculating decibel levels against a defined volume threshold to generate a list of timestamps for audible segments. These s

    Removes quiet sections from video recordings to create faster paced content without manual editing.

    Python
    Ver en GitHub↗3,141
  • zzw922cn/automatic_speech_recognitionAvatar de zzw922cn

    zzw922cn/Automatic_Speech_Recognition

    2,834Ver en GitHub↗

    Este proyecto es un kit de herramientas de machine learning diseñado para el desarrollo, entrenamiento y despliegue de motores de reconocimiento automático de voz. Proporciona un framework integral para convertir audio hablado en texto escrito, soportando específicamente modelos entrenados en datasets de mandarín e inglés. La librería utiliza una arquitectura neuronal end-to-end que procesa la entrada de audio cruda directamente en secuencias de caracteres, evitando la necesidad de una alineación lingüística intermedia. Incorpora técnicas de procesamiento de señales para transformar ondas sonoras en espectrogramas numéricos y vectores de características, que luego se utilizan para entrenar modelos acústicos a través de ciclos de aprendizaje iterativos acelerados por hardware. El kit de herramientas incluye un conjunto completo de utilidades para gestionar el ciclo de vida del modelo, incluyendo preprocesamiento de datos, persistencia de estado basada en puntos de control y evaluación de rendimiento. Los usuarios pueden evaluar la calidad de la transcripción calculando métricas como la distancia de edición de fonemas frente a etiquetas de verdad fundamental para cuantificar la precisión de la conversión de voz a texto.

    Standardizes and cleans raw audio data into numerical feature vectors for machine learning analysis.

    Pythonaudioautomatic-speech-recognitionchinese-speech-recognition
    Ver en GitHub↗2,834
  1. Home
  2. Artificial Intelligence & ML
  3. Dataset Preprocessing Tools
  4. Audio Dataset Preprocessing

Explorar subetiquetas

  • Speech-Based Silence RemovalAutomated tools that isolate speech segments and merge them to remove non-speech intervals. **Distinct from Audio Dataset Preprocessing:** Distinct from Audio Dataset Preprocessing: focuses on the functional removal of silence for downstream transcription rather than general dataset standardization.
  • Video Silence RemovalAutomated removal of silent sections from video files to increase pacing. **Distinct from Speech-Based Silence Removal:** Distinct from speech-based silence removal which targets audio datasets for transcription; this targets video pacing.