awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

13 रिपॉजिटरी

Awesome GitHub RepositoriesAudio Dataset Preprocessing

Tools for cleaning and standardizing raw audio and open-source speech datasets for ML training.

Distinct from Dataset Preprocessing Tools: Focuses on audio-specific cleaning and unification, whereas the parent is a general ML preprocessing utility.

Explore 13 awesome GitHub repositories matching artificial intelligence & ml · Audio Dataset Preprocessing. Refine with filters or upvote what's useful.

Awesome Audio Dataset Preprocessing GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • espnet/espnetespnet का अवतार

    espnet/espnet

    9,861GitHub पर देखें↗

    ESPnet is a comprehensive speech processing toolkit and PyTorch-based trainer designed for building end-to-end speech recognition, synthesis, and translation models. It provides a structured framework for developing automatic speech recognition systems using transducer and encoder-decoder architectures, alongside engines for text-to-speech synthesis and speech translation pipelines. The project distinguishes itself through a recipe-based workflow execution system that ensures experimental reproducibility by running standardized sequences of scripts for data preparation and model training. It

    Converts raw audio files into structured manifests required for model training and evaluation.

    Python
    GitHub पर देखें↗9,861
  • open-mmlab/amphionopen-mmlab का अवतार

    open-mmlab/Amphion

    9,844GitHub पर देखें↗

    Amphion is an audio generation toolkit designed for the research and development of models that synthesize speech, music, and environmental sound effects. It provides a standardized framework for reproducible audio synthesis, incorporating a text-to-speech engine and a voice conversion framework. The project specializes in transforming audio identities, allowing for the modification of speaker accents and voice identities while preserving original rhythm and style. It also includes capabilities for singing voice synthesis and the generation of environmental soundscapes from text descriptions

    Unifies the cleaning and preparation of various open-source audio datasets and raw speech data.

    Pythonaudio-generationaudio-synthesisaudioldm
    GitHub पर देखें↗9,844
  • voicepaw/so-vits-svc-forkvoicepaw का अवतार

    voicepaw/so-vits-svc-fork

    9,318GitHub पर देखें↗

    This project is an AI singing voice conversion system and vocal processor used for training generative voice models and converting vocal recordings or live input into a target voice. It functions as a VITS model trainer and a real-time voice changer that transforms vocal timbre and pitch to change the identity of a singer. The system provides a graphical management dashboard for controlling training hyperparameters and voice conversion presets. It supports low-latency audio streaming for live microphone input and employs pitch estimation to ensure precise matching between source and target vo

    Provides tools for cleaning, segmenting, and standardizing raw audio recordings for ML training.

    Pythoncontentvecdeep-learninggan
    GitHub पर देखें↗9,318
  • dusty-nv/jetson-inferencedusty-nv का अवतार

    dusty-nv/jetson-inference

    8,734GitHub पर देखें↗

    jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput. The project distinguishes itself through high-performance streaming analytics and the ability to execute concurrent AI pipelines on auto-grade silicon. It provides specialized support for multi-sensor stream processing, utilizing zero-copy data transport to load camera frames directly into GPU memory. The codebase covers a broad surface of capabiliti

    Transcribes and filters speech data using automatic speech recognition to prepare high-quality audio datasets.

    C++caffecomputer-visiondeep-learning
    GitHub पर देखें↗8,734
  • nl8590687/asrt_speechrecognitionnl8590687 का अवतार

    nl8590687/ASRT_SpeechRecognition

    8,375GitHub पर देखें↗

    This project is a Chinese automatic speech recognition framework and deep learning system designed to convert spoken Chinese audio into written text. It functions as a toolkit for training, evaluating, and deploying speech-to-text models, utilizing a specialized pinyin-to-text converter that transforms phonetic sequences into Chinese characters using a probability graph model. The system is distinguished by its deployment flexibility, offering a dockerized recognition server that provides transcription capabilities as a remote API. It supports high-performance streaming through a gRPC speech-

    Implements tools for cleaning and standardizing raw audio datasets specifically for machine learning training.

    Pythonasrtchinese-speech-recognitioncnn
    GitHub पर देखें↗8,375
  • snakers4/silero-vadsnakers4 का अवतार

    snakers4/silero-vad

    8,209GitHub पर देखें↗

    Silero VAD is a voice activity detection model and deep learning speech classifier designed to distinguish human speech from silence across diverse languages and noisy environments. It functions as a pre-trained neural network capable of identifying speech segments within both static audio recordings and real-time data streams. The project includes a language identification tool for classifying spoken languages and a framework for fine-tuning audio models. It provides utilities for optimizing detection thresholds using validation datasets and retraining the model with custom labeled audio to

    Isolates and merges speech segments from a recording to remove silence before transcription.

    Pythononnxonnx-runtimeonnxruntime
    GitHub पर देखें↗8,209
  • facebookresearch/auglyfacebookresearch का अवतार

    facebookresearch/AugLy

    5,086GitHub पर देखें↗

    AugLy एक मल्टीमॉडल डेटा ऑगमेंटेशन लाइब्रेरी और मशीन लर्निंग डेटासेट ऑगमेंटोर है। यह मॉडल की मजबूती (robustness) को बेहतर बनाने और सैंपल विविधता को बढ़ाने के लिए ऑडियो, इमेज, टेक्स्ट और वीडियो डेटासेट में ट्रेनिंग डेटा के सिंथेटिक वेरिएशन उत्पन्न करने के लिए एक सिस्टम प्रदान करती है। यह लाइब्रेरी एक मल्टीमीडिया नॉइज़ सिम्युलेटर के रूप में कार्य करती है, जिसे विशेष रूप से सोशल मीडिया टेम्पलेट्स और इंटरनेट आर्टिफैक्ट्स को मीडिया पर ओवरले करके वास्तविक दुनिया के यूजर कैप्चर की नकल करने के लिए डिज़ाइन किया गया है। इसमें प्रत्येक ऑगमेंटेड डेटा के टुकड़े पर लागू किए गए विशिष्ट ट्रांसफ़ॉर्मेशन्स और तीव्रता स्तरों को रिकॉर्ड करने के लिए एक डेटा प्रोवेनेंस ट्रैकर शामिल है। यह टूल टेक्स्ट के लिए भाषाई ट्रांसफ़ॉर्मेशन्स, वीडियो के लिए टेम्पोरल और विज़ुअल ट्रांसफ़ॉर्मेशन्स, और ऑडियो के लिए सोनिक ट्रांसफ़ॉर्मेशन्स सहित डेटासेट विस्तार क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है।

    Applies transformations to audio files to create more varied training samples for sound recognition or processing models.

    Python
    GitHub पर देखें↗5,086
  • microsoft/muzicmicrosoft का अवतार

    microsoft/muzic

    4,928GitHub पर देखें↗

    Muzic AI-संचालित संगीत विश्लेषण, रचना और संश्लेषण के लिए एक डीप लर्निंग प्लेटफ़ॉर्म और फ्रेमवर्क है। यह एक संगीत जनरेशन फ्रेमवर्क और विश्लेषण टूल के रूप में कार्य करता है, जो प्रतीकात्मक और ऑडियो संगीत के निर्माण और व्याख्या को व्यवस्थित करने के लिए बड़े भाषा मॉडल्स और स्वायत्त एजेंटों का उपयोग करता है। यह प्रोजेक्ट अपनी क्रॉस-मॉडल क्षमताओं द्वारा प्रतिष्ठित है, जो ज़ीरो-शॉट वर्गीकरण और सूचना पुनर्प्राप्ति के लिए प्राकृतिक भाषा और प्रतीकात्मक संगीत को एक साझा संयुक्त एम्बेडिंग स्पेस में मैप करता है। यह विभिन्न प्रकार के विशेष आर्किटेक्चर को नियोजित करता है, जिसमें ऑडियो संश्लेषण के लिए डिफ्यूज़न फ्रेमवर्क, लंबी-अनुक्रम संरचनात्मक स्थिरता के लिए डुअल-ग्रेन अटेंशन मैकेनिज्म और एक हाइब्रिड सिस्टम शामिल है जो न्यूरल नेटवर्क के साथ संगीत सिद्धांत नियमों को जोड़ता है। यह प्लेटफ़ॉर्म टेक्स्ट और लिरिक्स से MIDI अनुक्रमों के निर्माण, न्यूरल सिंगिंग वॉयस सिंथेसिस और स्वचालित लिरिक्स ट्रांसक्रिप्शन सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह संगीत संरचना मॉडलिंग, विशेषता-आधारित प्रतीकात्मक जनरेशन और स्वायत्त एजेंटों के माध्यम से बाहरी संगीत टूल्स के ऑर्केस्ट्रेशन के लिए टूल्स भी प्रदान करता है। सहायक यूटिलिटीज में बड़े पैमाने पर MIDI बाइनराइजेशन, डेटासेट एन्कोडिंग और मेलोडी नोट निष्कर्षण और स्पीच-टू-फोनम एलाइनमेंट के लिए ऑडियो सिग्नल प्रोसेसिंग के लिए डेटा इंजीनियरिंग पाइपलाइन शामिल हैं।

    Provides tools for cleaning and converting raw MIDI and audio files into formats suitable for ML training.

    Pythonai-musicdeep-learningmusic
    GitHub पर देखें↗4,928
  • buriburisuri/speech-to-text-wavenetburiburisuri का अवतार

    buriburisuri/speech-to-text-wavenet

    4,007GitHub पर देखें↗

    This project is a deep learning framework designed for end-to-end speech-to-text transcription. It utilizes the WaveNet neural network architecture to process spoken audio input and generate written text transcripts, leveraging connectionist temporal classification to map variable-length audio sequences to character-level outputs. The system distinguishes itself through a comprehensive training pipeline that supports distributed execution across multiple graphics processing units. It includes specialized utilities for audio data augmentation and the transformation of raw audio files into opti

    Transforms raw audio files into optimized feature formats to accelerate machine learning training and reduce disk input bottlenecks.

    Python
    GitHub पर देखें↗4,007
  • tensorspeech/tensorflowttsTensorSpeech का अवतार

    TensorSpeech/TensorflowTTS

    3,993GitHub पर देखें↗

    TensorFlowTTS is a neural speech synthesis framework used to convert text into high-fidelity audio waveforms. It provides a toolkit for training and fine-tuning sequence-to-sequence or generative adversarial network architectures to produce natural sounding speech. The system includes neural vocoder implementations that transform intermediate acoustic representations into final audio waveforms. It also features playback speed control to adjust the rate of synthesized speech output. The framework covers the end-to-end pipeline for speech synthesis, including audio data preprocessing to create

    Provides utilities to convert raw audio and transcriptions into normalized mel spectrograms for ML training.

    Python
    GitHub पर देखें↗3,993
  • stability-ai/stable-audio-toolsStability-AI का अवतार

    Stability-AI/stable-audio-tools

    3,790GitHub पर देखें↗

    Stable-audio-tools is a toolkit for training and deploying latent diffusion models for high-fidelity audio synthesis. It provides a framework for generating audio by iteratively refining noise within a compressed latent space, using specialized encoders to preserve temporal and spectral features of the audio signal. The project features a system for adapting pre-trained audio checkpoints to new datasets through modular initialization and configuration files. It includes utilities for weight extraction and inference model export, which remove training metadata and optimizer states to create li

    Integrates audio data from local directories or cloud stores for use in machine learning training pipelines.

    Python
    GitHub पर देखें↗3,790
  • carykh/jumpcuttercarykh का अवतार

    carykh/jumpcutter

    3,141GitHub पर देखें↗

    Jumpcutter is an audio-based video cutter and automatic editor designed to eliminate dead air from video files. It functions as a utility that condenses footage by detecting and removing silent sections based on audio track analysis. The tool utilizes FFmpeg to automatically identify quiet gaps and strip them from recordings. This process focuses on removing silent video sections to create faster-paced content without the need for manual editing. The system operates by calculating decibel levels against a defined volume threshold to generate a list of timestamps for audible segments. These s

    Removes quiet sections from video recordings to create faster paced content without manual editing.

    Python
    GitHub पर देखें↗3,141
  • zzw922cn/automatic_speech_recognitionzzw922cn का अवतार

    zzw922cn/Automatic_Speech_Recognition

    2,834GitHub पर देखें↗

    This project is a machine learning toolkit designed for the development, training, and deployment of automatic speech recognition engines. It provides a comprehensive framework for converting spoken audio into written text, specifically supporting models trained on Mandarin and English datasets. The library utilizes an end-to-end neural architecture that processes raw audio input directly into character sequences, bypassing the need for intermediate linguistic alignment. It incorporates signal processing techniques to transform sound waves into numerical spectrograms and feature vectors, whic

    Standardizes and cleans raw audio data into numerical feature vectors for machine learning analysis.

    Pythonaudioautomatic-speech-recognitionchinese-speech-recognition
    GitHub पर देखें↗2,834
  1. Home
  2. Artificial Intelligence & ML
  3. Dataset Preprocessing Tools
  4. Audio Dataset Preprocessing

सब-टैग एक्सप्लोर करें

  • Speech-Based Silence RemovalAutomated tools that isolate speech segments and merge them to remove non-speech intervals. **Distinct from Audio Dataset Preprocessing:** Distinct from Audio Dataset Preprocessing: focuses on the functional removal of silence for downstream transcription rather than general dataset standardization.
  • Video Silence RemovalAutomated removal of silent sections from video files to increase pacing. **Distinct from Speech-Based Silence Removal:** Distinct from speech-based silence removal which targets audio datasets for transcription; this targets video pacing.