13 repositorios
Tools for cleaning and standardizing raw audio and open-source speech datasets for ML training.
Distinct from Dataset Preprocessing Tools: Focuses on audio-specific cleaning and unification, whereas the parent is a general ML preprocessing utility.
Explore 13 awesome GitHub repositories matching artificial intelligence & ml · Audio Dataset Preprocessing. Refine with filters or upvote what's useful.
ESPnet is a comprehensive speech processing toolkit and PyTorch-based trainer designed for building end-to-end speech recognition, synthesis, and translation models. It provides a structured framework for developing automatic speech recognition systems using transducer and encoder-decoder architectures, alongside engines for text-to-speech synthesis and speech translation pipelines. The project distinguishes itself through a recipe-based workflow execution system that ensures experimental reproducibility by running standardized sequences of scripts for data preparation and model training. It
Converts raw audio files into structured manifests required for model training and evaluation.
Amphion is an audio generation toolkit designed for the research and development of models that synthesize speech, music, and environmental sound effects. It provides a standardized framework for reproducible audio synthesis, incorporating a text-to-speech engine and a voice conversion framework. The project specializes in transforming audio identities, allowing for the modification of speaker accents and voice identities while preserving original rhythm and style. It also includes capabilities for singing voice synthesis and the generation of environmental soundscapes from text descriptions
Unifies the cleaning and preparation of various open-source audio datasets and raw speech data.
This project is an AI singing voice conversion system and vocal processor used for training generative voice models and converting vocal recordings or live input into a target voice. It functions as a VITS model trainer and a real-time voice changer that transforms vocal timbre and pitch to change the identity of a singer. The system provides a graphical management dashboard for controlling training hyperparameters and voice conversion presets. It supports low-latency audio streaming for live microphone input and employs pitch estimation to ensure precise matching between source and target vo
Provides tools for cleaning, segmenting, and standardizing raw audio recordings for ML training.
jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput. The project distinguishes itself through high-performance streaming analytics and the ability to execute concurrent AI pipelines on auto-grade silicon. It provides specialized support for multi-sensor stream processing, utilizing zero-copy data transport to load camera frames directly into GPU memory. The codebase covers a broad surface of capabiliti
Transcribes and filters speech data using automatic speech recognition to prepare high-quality audio datasets.
This project is a Chinese automatic speech recognition framework and deep learning system designed to convert spoken Chinese audio into written text. It functions as a toolkit for training, evaluating, and deploying speech-to-text models, utilizing a specialized pinyin-to-text converter that transforms phonetic sequences into Chinese characters using a probability graph model. The system is distinguished by its deployment flexibility, offering a dockerized recognition server that provides transcription capabilities as a remote API. It supports high-performance streaming through a gRPC speech-
Implements tools for cleaning and standardizing raw audio datasets specifically for machine learning training.
Silero VAD is a voice activity detection model and deep learning speech classifier designed to distinguish human speech from silence across diverse languages and noisy environments. It functions as a pre-trained neural network capable of identifying speech segments within both static audio recordings and real-time data streams. The project includes a language identification tool for classifying spoken languages and a framework for fine-tuning audio models. It provides utilities for optimizing detection thresholds using validation datasets and retraining the model with custom labeled audio to
Isolates and merges speech segments from a recording to remove silence before transcription.
AugLy es una biblioteca de aumento de datos multimodal y aumentador de conjuntos de datos de machine learning. Proporciona un sistema para generar variaciones sintéticas de datos de entrenamiento a través de conjuntos de datos de audio, imagen, texto y video para aumentar la diversidad de muestras y mejorar la robustez del modelo. La biblioteca funciona como un simulador de ruido multimedia, diseñado específicamente para imitar capturas de usuarios del mundo real superponiendo plantillas de redes sociales y artefactos de internet sobre los medios. Incluye un rastreador de procedencia de datos para registrar las transformaciones específicas y los niveles de intensidad aplicados a cada pieza de datos aumentados. La herramienta cubre una amplia gama de capacidades de expansión de conjuntos de datos, incluyendo transformaciones lingüísticas para texto, transformaciones temporales y visuales para video, y transformaciones sónicas para audio.
Applies transformations to audio files to create more varied training samples for sound recognition or processing models.
Muzic es una plataforma y framework de deep learning para el análisis, composición y síntesis de música impulsada por IA. Funciona como un framework de generación de música y herramienta de análisis, utilizando modelos de lenguaje grandes y agentes autónomos para orquestar la creación e interpretación de música simbólica y de audio. El proyecto se distingue por sus capacidades multimodales, mapeando el lenguaje natural y la música simbólica en un espacio de incrustación (embedding) conjunto compartido para clasificación zero-shot y recuperación de información. Emplea una variedad de arquitecturas especializadas, incluyendo frameworks de difusión para síntesis de audio, mecanismos de atención de grano dual para consistencia estructural de secuencias largas y un sistema híbrido que combina reglas de teoría musical con redes neuronales. La plataforma cubre una amplia gama de capacidades, incluyendo la generación de secuencias MIDI a partir de texto y letras, síntesis de voz cantada neuronal y transcripción automatizada de letras. También proporciona herramientas para el modelado de estructuras musicales, generación simbólica basada en atributos y la orquestación de herramientas musicales externas a través de agentes autónomos. Las utilidades de soporte incluyen pipelines de ingeniería de datos para la binarización de MIDI a gran escala, codificación de conjuntos de datos y procesamiento de señales de audio para la extracción de notas de melodía y alineación de voz a fonema.
Provides tools for cleaning and converting raw MIDI and audio files into formats suitable for ML training.
Este proyecto es un framework de aprendizaje profundo diseñado para la transcripción de voz a texto de extremo a extremo. Utiliza la arquitectura de red neuronal WaveNet para procesar la entrada de audio hablado y generar transcripciones de texto escrito, aprovechando la clasificación temporal conexionista (CTC) para mapear secuencias de audio de longitud variable a salidas a nivel de carácter. El sistema se distingue por una tubería de entrenamiento integral que soporta la ejecución distribuida a través de múltiples unidades de procesamiento gráfico (GPU). Incluye utilidades especializadas para la aumentación de datos de audio y la transformación de archivos de audio crudos en formatos binarios optimizados, lo que minimiza la latencia de entrada y salida de disco durante el entrenamiento de modelos a gran escala. El software proporciona un entorno completo para gestionar flujos de trabajo de aprendizaje automático, incluyendo herramientas para calcular métricas de pérdida para monitorear la convergencia y precisión del modelo. Todos los componentes, incluyendo el motor de reconocimiento y las tuberías de entrenamiento, están diseñados para el despliegue dentro de entornos contenedorizados para asegurar una ejecución consistente en diversos sistemas host.
Transforms raw audio files into optimized feature formats to accelerate machine learning training and reduce disk input bottlenecks.
TensorFlowTTS is a neural speech synthesis framework used to convert text into high-fidelity audio waveforms. It provides a toolkit for training and fine-tuning sequence-to-sequence or generative adversarial network architectures to produce natural sounding speech. The system includes neural vocoder implementations that transform intermediate acoustic representations into final audio waveforms. It also features playback speed control to adjust the rate of synthesized speech output. The framework covers the end-to-end pipeline for speech synthesis, including audio data preprocessing to create
Provides utilities to convert raw audio and transcriptions into normalized mel spectrograms for ML training.
Stable-audio-tools is a toolkit for training and deploying latent diffusion models for high-fidelity audio synthesis. It provides a framework for generating audio by iteratively refining noise within a compressed latent space, using specialized encoders to preserve temporal and spectral features of the audio signal. The project features a system for adapting pre-trained audio checkpoints to new datasets through modular initialization and configuration files. It includes utilities for weight extraction and inference model export, which remove training metadata and optimizer states to create li
Integrates audio data from local directories or cloud stores for use in machine learning training pipelines.
Jumpcutter is an audio-based video cutter and automatic editor designed to eliminate dead air from video files. It functions as a utility that condenses footage by detecting and removing silent sections based on audio track analysis. The tool utilizes FFmpeg to automatically identify quiet gaps and strip them from recordings. This process focuses on removing silent video sections to create faster-paced content without the need for manual editing. The system operates by calculating decibel levels against a defined volume threshold to generate a list of timestamps for audible segments. These s
Removes quiet sections from video recordings to create faster paced content without manual editing.
Este proyecto es un kit de herramientas de machine learning diseñado para el desarrollo, entrenamiento y despliegue de motores de reconocimiento automático de voz. Proporciona un framework integral para convertir audio hablado en texto escrito, soportando específicamente modelos entrenados en datasets de mandarín e inglés. La librería utiliza una arquitectura neuronal end-to-end que procesa la entrada de audio cruda directamente en secuencias de caracteres, evitando la necesidad de una alineación lingüística intermedia. Incorpora técnicas de procesamiento de señales para transformar ondas sonoras en espectrogramas numéricos y vectores de características, que luego se utilizan para entrenar modelos acústicos a través de ciclos de aprendizaje iterativos acelerados por hardware. El kit de herramientas incluye un conjunto completo de utilidades para gestionar el ciclo de vida del modelo, incluyendo preprocesamiento de datos, persistencia de estado basada en puntos de control y evaluación de rendimiento. Los usuarios pueden evaluar la calidad de la transcripción calculando métricas como la distancia de edición de fonemas frente a etiquetas de verdad fundamental para cuantificar la precisión de la conversión de voz a texto.
Standardizes and cleans raw audio data into numerical feature vectors for machine learning analysis.