awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

13 个仓库

Awesome GitHub RepositoriesAudio Dataset Preprocessing

Tools for cleaning and standardizing raw audio and open-source speech datasets for ML training.

Distinct from Dataset Preprocessing Tools: Focuses on audio-specific cleaning and unification, whereas the parent is a general ML preprocessing utility.

Explore 13 awesome GitHub repositories matching artificial intelligence & ml · Audio Dataset Preprocessing. Refine with filters or upvote what's useful.

Awesome Audio Dataset Preprocessing GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • espnet/espnetespnet 的头像

    espnet/espnet

    9,861在 GitHub 上查看↗

    ESPnet is a comprehensive speech processing toolkit and PyTorch-based trainer designed for building end-to-end speech recognition, synthesis, and translation models. It provides a structured framework for developing automatic speech recognition systems using transducer and encoder-decoder architectures, alongside engines for text-to-speech synthesis and speech translation pipelines. The project distinguishes itself through a recipe-based workflow execution system that ensures experimental reproducibility by running standardized sequences of scripts for data preparation and model training. It

    Converts raw audio files into structured manifests required for model training and evaluation.

    Python
    在 GitHub 上查看↗9,861
  • open-mmlab/amphionopen-mmlab 的头像

    open-mmlab/Amphion

    9,844在 GitHub 上查看↗

    Amphion is an audio generation toolkit designed for the research and development of models that synthesize speech, music, and environmental sound effects. It provides a standardized framework for reproducible audio synthesis, incorporating a text-to-speech engine and a voice conversion framework. The project specializes in transforming audio identities, allowing for the modification of speaker accents and voice identities while preserving original rhythm and style. It also includes capabilities for singing voice synthesis and the generation of environmental soundscapes from text descriptions

    Unifies the cleaning and preparation of various open-source audio datasets and raw speech data.

    Pythonaudio-generationaudio-synthesisaudioldm
    在 GitHub 上查看↗9,844
  • voicepaw/so-vits-svc-forkvoicepaw 的头像

    voicepaw/so-vits-svc-fork

    9,318在 GitHub 上查看↗

    This project is an AI singing voice conversion system and vocal processor used for training generative voice models and converting vocal recordings or live input into a target voice. It functions as a VITS model trainer and a real-time voice changer that transforms vocal timbre and pitch to change the identity of a singer. The system provides a graphical management dashboard for controlling training hyperparameters and voice conversion presets. It supports low-latency audio streaming for live microphone input and employs pitch estimation to ensure precise matching between source and target vo

    Provides tools for cleaning, segmenting, and standardizing raw audio recordings for ML training.

    Pythoncontentvecdeep-learninggan
    在 GitHub 上查看↗9,318
  • dusty-nv/jetson-inferencedusty-nv 的头像

    dusty-nv/jetson-inference

    8,734在 GitHub 上查看↗

    jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput. The project distinguishes itself through high-performance streaming analytics and the ability to execute concurrent AI pipelines on auto-grade silicon. It provides specialized support for multi-sensor stream processing, utilizing zero-copy data transport to load camera frames directly into GPU memory. The codebase covers a broad surface of capabiliti

    Transcribes and filters speech data using automatic speech recognition to prepare high-quality audio datasets.

    C++caffecomputer-visiondeep-learning
    在 GitHub 上查看↗8,734
  • nl8590687/asrt_speechrecognitionnl8590687 的头像

    nl8590687/ASRT_SpeechRecognition

    8,375在 GitHub 上查看↗

    This project is a Chinese automatic speech recognition framework and deep learning system designed to convert spoken Chinese audio into written text. It functions as a toolkit for training, evaluating, and deploying speech-to-text models, utilizing a specialized pinyin-to-text converter that transforms phonetic sequences into Chinese characters using a probability graph model. The system is distinguished by its deployment flexibility, offering a dockerized recognition server that provides transcription capabilities as a remote API. It supports high-performance streaming through a gRPC speech-

    Implements tools for cleaning and standardizing raw audio datasets specifically for machine learning training.

    Pythonasrtchinese-speech-recognitioncnn
    在 GitHub 上查看↗8,375
  • snakers4/silero-vadsnakers4 的头像

    snakers4/silero-vad

    8,209在 GitHub 上查看↗

    Silero VAD is a voice activity detection model and deep learning speech classifier designed to distinguish human speech from silence across diverse languages and noisy environments. It functions as a pre-trained neural network capable of identifying speech segments within both static audio recordings and real-time data streams. The project includes a language identification tool for classifying spoken languages and a framework for fine-tuning audio models. It provides utilities for optimizing detection thresholds using validation datasets and retraining the model with custom labeled audio to

    Isolates and merges speech segments from a recording to remove silence before transcription.

    Pythononnxonnx-runtimeonnxruntime
    在 GitHub 上查看↗8,209
  • facebookresearch/auglyfacebookresearch 的头像

    facebookresearch/AugLy

    5,086在 GitHub 上查看↗

    AugLy 是一个多模态数据增强库和机器学习数据集增强器。它提供了一个系统,用于在音频、图像、文本和视频数据集上生成训练数据的合成变体,以增加样本多样性并提高模型鲁棒性。 该库作为一个多媒体噪声模拟器,专门设计用于通过在媒体上叠加社交媒体模板和互联网伪影来模拟真实世界的用户捕获。它包括一个数据来源跟踪器,用于记录应用于每条增强数据的特定转换和强度级别。 该工具涵盖了广泛的数据集扩展功能,包括文本的语言转换、视频的时间和视觉转换以及音频的声学转换。

    Applies transformations to audio files to create more varied training samples for sound recognition or processing models.

    Python
    在 GitHub 上查看↗5,086
  • microsoft/muzicmicrosoft 的头像

    microsoft/muzic

    4,928在 GitHub 上查看↗

    Muzic 是一个用于 AI 驱动的音乐分析、创作和合成的深度学习平台和框架。它作为一个音乐生成框架和分析工具,利用大型语言模型和自主智能体来编排符号音乐和音频音乐的创作与解读。 该项目以其跨模态能力而著称,将自然语言和符号音乐映射到共享的联合嵌入空间中,用于零样本分类和信息检索。它采用了多种专门的架构,包括用于音频合成的扩散框架、用于长序列结构一致性的双粒度注意力机制,以及结合音乐理论规则与神经网络的混合系统。 该平台涵盖了广泛的功能,包括从文本和歌词生成 MIDI 序列、神经歌声合成以及自动歌词转录。它还提供用于音乐结构建模、基于属性的符号生成以及通过自主智能体编排外部音乐工具的工具。 支持性实用程序包括用于大规模 MIDI 二进制化、数据集编码的数据工程流水线,以及用于旋律音符提取和语音到音素对齐的音频信号处理。

    Provides tools for cleaning and converting raw MIDI and audio files into formats suitable for ML training.

    Pythonai-musicdeep-learningmusic
    在 GitHub 上查看↗4,928
  • buriburisuri/speech-to-text-wavenetburiburisuri 的头像

    buriburisuri/speech-to-text-wavenet

    4,007在 GitHub 上查看↗

    这是一个专为端到端语音转文字转录设计的深度学习框架。它利用 WaveNet 神经网络架构处理语音音频输入并生成文本转录,通过联结主义时间分类(CTC)将变长音频序列映射为字符级输出。 该系统的特色在于其全面的训练流水线,支持跨多个 GPU 的分布式执行。它包含用于音频数据增强的专用工具,以及将原始音频文件转换为优化二进制格式的实用程序,从而最大限度地减少大规模模型训练期间的磁盘 I/O 延迟。 该软件为管理机器学习工作流提供了完整环境,包括用于计算损失指标以监控模型收敛性和准确性的工具。包括识别引擎和训练流水线在内的所有组件,均设计为可在容器化环境中部署,以确保在不同宿主系统上的一致性执行。

    Transforms raw audio files into optimized feature formats to accelerate machine learning training and reduce disk input bottlenecks.

    Python
    在 GitHub 上查看↗4,007
  • tensorspeech/tensorflowttsTensorSpeech 的头像

    TensorSpeech/TensorflowTTS

    3,993在 GitHub 上查看↗

    TensorFlowTTS 是一个神经语音合成框架,用于将文本转换为高保真音频波形。它提供了一个用于训练和微调序列到序列或生成对抗网络架构的工具包,以产生自然听感的语音。 该系统包括将中间声学表示转换为最终音频波形的神经声码器实现。它还具有播放速度控制功能,以调整合成语音输出的速率。 该框架涵盖了语音合成的端到端流水线,包括用于创建归一化梅尔频谱图的音频数据预处理,以及用于管理 GPU 加速模型训练的训练流水线。它利用自定义训练器框架在训练过程中处理损失函数和优化逻辑。

    Provides utilities to convert raw audio and transcriptions into normalized mel spectrograms for ML training.

    Python
    在 GitHub 上查看↗3,993
  • stability-ai/stable-audio-toolsStability-AI 的头像

    Stability-AI/stable-audio-tools

    3,790在 GitHub 上查看↗

    Stable-audio-tools is a toolkit for training and deploying latent diffusion models for high-fidelity audio synthesis. It provides a framework for generating audio by iteratively refining noise within a compressed latent space, using specialized encoders to preserve temporal and spectral features of the audio signal. The project features a system for adapting pre-trained audio checkpoints to new datasets through modular initialization and configuration files. It includes utilities for weight extraction and inference model export, which remove training metadata and optimizer states to create li

    Integrates audio data from local directories or cloud stores for use in machine learning training pipelines.

    Python
    在 GitHub 上查看↗3,790
  • carykh/jumpcuttercarykh 的头像

    carykh/jumpcutter

    3,141在 GitHub 上查看↗

    Jumpcutter is an audio-based video cutter and automatic editor designed to eliminate dead air from video files. It functions as a utility that condenses footage by detecting and removing silent sections based on audio track analysis. The tool utilizes FFmpeg to automatically identify quiet gaps and strip them from recordings. This process focuses on removing silent video sections to create faster-paced content without the need for manual editing. The system operates by calculating decibel levels against a defined volume threshold to generate a list of timestamps for audible segments. These s

    Removes quiet sections from video recordings to create faster paced content without manual editing.

    Python
    在 GitHub 上查看↗3,141
  • zzw922cn/automatic_speech_recognitionzzw922cn 的头像

    zzw922cn/Automatic_Speech_Recognition

    2,834在 GitHub 上查看↗

    该项目是一个机器学习工具包,专为自动语音识别引擎的开发、训练和部署而设计。它提供了一个将口语音频转换为书面文本的综合框架,特别支持在中文和英文数据集上训练的模型。 该库利用端到端的神经架构,直接将原始音频输入处理为字符序列,无需中间的语言对齐。它结合了信号处理技术,将声波转换为数值频谱图和特征向量,然后通过迭代的、硬件加速的学习周期来训练声学模型。 该工具包包含一套完整的模型生命周期管理实用程序,包括数据预处理、基于检查点的状态持久化以及性能评估。用户可以通过计算诸如音素编辑距离等指标来评估转录质量,从而量化语音转文本转换的精度。

    Standardizes and cleans raw audio data into numerical feature vectors for machine learning analysis.

    Pythonaudioautomatic-speech-recognitionchinese-speech-recognition
    在 GitHub 上查看↗2,834
  1. Home
  2. Artificial Intelligence & ML
  3. Dataset Preprocessing Tools
  4. Audio Dataset Preprocessing

探索子标签

  • Speech-Based Silence RemovalAutomated tools that isolate speech segments and merge them to remove non-speech intervals. **Distinct from Audio Dataset Preprocessing:** Distinct from Audio Dataset Preprocessing: focuses on the functional removal of silence for downstream transcription rather than general dataset standardization.
  • Video Silence RemovalAutomated removal of silent sections from video files to increase pacing. **Distinct from Speech-Based Silence Removal:** Distinct from speech-based silence removal which targets audio datasets for transcription; this targets video pacing.