awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

14 रिपॉजिटरी

Awesome GitHub RepositoriesDeep Learning Audio Libraries

Libraries focused on high-fidelity audio synthesis and processing using neural architectures.

Distinct from Deep Learning Libraries: Shortlist targets general DL libraries or simple audio processing; this is a deep learning library for synthesis.

Explore 14 awesome GitHub repositories matching artificial intelligence & ml · Deep Learning Audio Libraries. Refine with filters or upvote what's useful.

Awesome Deep Learning Audio Libraries GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • deezer/spleeterdeezer का अवतार

    deezer/spleeter

    28,252GitHub पर देखें↗

    Spleeter is an AI audio source separation library and deep learning toolkit designed to split mixed music files into individual audio stems, such as vocals and drums. It provides a suite of pretrained models for isolating different instruments and voices from a recording. The toolkit includes capabilities for training and evaluating custom audio separation models using labeled datasets and configuration files. It also features utilities for measuring model performance by comparing separation outputs against reference datasets. The system manages audio processing through spectral representati

    Functions as a deep learning library for splitting music files into individual audio stems.

    Pythonaudio-processingbassdeep-learning
    GitHub पर देखें↗28,252
  • facebookresearch/audiocraftfacebookresearch का अवतार

    facebookresearch/audiocraft

    23,379GitHub पर देखें↗

    Audiocraft is a deep learning audio library and machine learning framework designed for training, fine-tuning, and evaluating generative models for music and sound effects. It functions as a text-to-music generative model and a neural audio codec, providing the tools necessary to compress audio signals into discrete representations and synthesize high-fidelity waveforms from textual descriptions. The framework is distinguished by its ability to combine multiple conditioning signals, allowing for the generation of audio based on text prompts, melodic excerpts, or style-based audio clips. It al

    Ships a library for processing and generating high-fidelity audio using neural networks and transformer models.

    Jupyter Notebook
    GitHub पर देखें↗23,379
  • wwmm/easyeffectswwmm का अवतार

    wwmm/easyeffects

    9,690GitHub पर देखें↗

    EasyEffects is a real-time audio processor and system-wide effects manager designed for PipeWire audio streams. It functions as a comprehensive suite for applying filters, equalizers, and limiters to both input and output audio across the entire system. The project distinguishes itself through its use of deep learning for neural network noise suppression and voice isolation, as well as its ability to simulate physical acoustic environments using impulse-response convolution. It includes a sophisticated preset management system that allows users to associate specific audio configurations with

    Uses deep learning models to isolate voice from ambient noise and preserve speech intelligibility.

    HTMLauto-volumecompressorequalizer
    GitHub पर देखें↗9,690
  • jaywalnut310/vitsjaywalnut310 का अवतार

    jaywalnut310/vits

    7,862GitHub पर देखें↗

    This project is an end-to-end text-to-speech engine and deep learning voice synthesizer. It functions as a neural speech synthesis framework that converts written text directly into audio waveforms using a single neural network. The system implements an adversarial framework and a conditional variational autoencoder to generate high-fidelity artificial speech. It utilizes a generative adversarial network to ensure synthesized audio is indistinguishable from real human speech. The toolkit provides capabilities for neural speech synthesis, text-to-audio generation, and the training of custom v

    Functions as a deep learning audio library for training high-fidelity speech models from text and audio.

    Pythondeep-learningpytorchspeech-synthesis
    GitHub पर देखें↗7,862
  • ibab/tensorflow-wavenetibab का अवतार

    ibab/tensorflow-wavenet

    5,432GitHub पर देखें↗

    यह प्रोजेक्ट रॉ ऑडियो वेवफॉर्म जनरेशन के लिए एक न्यूरल नेटवर्क का TensorFlow कार्यान्वयन है। यह एक कंडीशन-आधारित स्पीच सिंथेसिस मॉडल के रूप में कार्य करता है जो डायलेटेड कन्वेन्शनल न्यूरल नेटवर्क आर्किटेक्चर का उपयोग करके सिंथेटिक ऑडियो सैंपल तैयार करता है। यह सिस्टम ट्रेनिंग और जनरेशन के दौरान ग्लोबल कंडीशनिंग और कैटेगोरिकल आइडेंटिफायर्स को शामिल करके कस्टम वॉयस मॉडलिंग का समर्थन करता है। यह मॉडल को न्यूरल टेक्स्ट-टू-स्पीच एप्लिकेशन के लिए विशिष्ट वक्ताओं या अलग-अलग ऑडियो विशेषताओं की नकल करने की अनुमति देता है। यह फ्रेमवर्क डीप लर्निंग ऑडियो सिंथेसिस को कवर करता है, जिसमें ऑडियो डेटासेट प्रोसेसिंग, वेवफॉर्म फाइलों से मॉडल ट्रेनिंग, और चलाने योग्य ऑडियो फाइलों का जनरेशन शामिल है। यह ऑडियो डेटा में लॉन्ग-रेंज डिपेंडेंसीज को संभालने के लिए डायलेटेड कॉज़ल कन्वेन्शन्स, mu-law कंपैंडिंग और क्वांटाइज्ड सॉफ्टमैक्स आउटपुट जैसे तकनीकी घटकों का उपयोग करता है।

    Implements a deep learning library for high-fidelity audio synthesis using dilated convolutional architectures.

    Python
    GitHub पर देखें↗5,432
  • microsoft/muzicmicrosoft का अवतार

    microsoft/muzic

    4,928GitHub पर देखें↗

    Muzic AI-संचालित संगीत विश्लेषण, रचना और संश्लेषण के लिए एक डीप लर्निंग प्लेटफ़ॉर्म और फ्रेमवर्क है। यह एक संगीत जनरेशन फ्रेमवर्क और विश्लेषण टूल के रूप में कार्य करता है, जो प्रतीकात्मक और ऑडियो संगीत के निर्माण और व्याख्या को व्यवस्थित करने के लिए बड़े भाषा मॉडल्स और स्वायत्त एजेंटों का उपयोग करता है। यह प्रोजेक्ट अपनी क्रॉस-मॉडल क्षमताओं द्वारा प्रतिष्ठित है, जो ज़ीरो-शॉट वर्गीकरण और सूचना पुनर्प्राप्ति के लिए प्राकृतिक भाषा और प्रतीकात्मक संगीत को एक साझा संयुक्त एम्बेडिंग स्पेस में मैप करता है। यह विभिन्न प्रकार के विशेष आर्किटेक्चर को नियोजित करता है, जिसमें ऑडियो संश्लेषण के लिए डिफ्यूज़न फ्रेमवर्क, लंबी-अनुक्रम संरचनात्मक स्थिरता के लिए डुअल-ग्रेन अटेंशन मैकेनिज्म और एक हाइब्रिड सिस्टम शामिल है जो न्यूरल नेटवर्क के साथ संगीत सिद्धांत नियमों को जोड़ता है। यह प्लेटफ़ॉर्म टेक्स्ट और लिरिक्स से MIDI अनुक्रमों के निर्माण, न्यूरल सिंगिंग वॉयस सिंथेसिस और स्वचालित लिरिक्स ट्रांसक्रिप्शन सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह संगीत संरचना मॉडलिंग, विशेषता-आधारित प्रतीकात्मक जनरेशन और स्वायत्त एजेंटों के माध्यम से बाहरी संगीत टूल्स के ऑर्केस्ट्रेशन के लिए टूल्स भी प्रदान करता है। सहायक यूटिलिटीज में बड़े पैमाने पर MIDI बाइनराइजेशन, डेटासेट एन्कोडिंग और मेलोडी नोट निष्कर्षण और स्पीच-टू-फोनम एलाइनमेंट के लिए ऑडियो सिग्नल प्रोसेसिंग के लिए डेटा इंजीनियरिंग पाइपलाइन शामिल हैं।

    Synthesizes musical audio and MIDI sequences using neural networks and deep learning models.

    Pythonai-musicdeep-learningmusic
    GitHub पर देखें↗4,928
  • google/lyragoogle का अवतार

    google/lyra

    3,964GitHub पर देखें↗

    Lyra is a voice compression framework and low-bitrate speech codec designed to transmit high-quality audio over bandwidth-constrained networks. It utilizes an adaptive bitrate audio codec to balance audio quality and network bandwidth during active sessions. The project employs generative audio compression, using neural networks to synthesize speech signals from minimal data and reconstruct missing audio details. This allows for high-quality voice audio reconstruction from highly compressed byte streams. The system covers bandwidth-optimized voice over IP and real-time voice communication, f

    Recreates high-quality speech signals from compressed byte streams using deep learning models.

    C++
    GitHub पर देखें↗3,964
  • andabi/deep-voice-conversionandabi का अवतार

    andabi/deep-voice-conversion

    3,941GitHub पर देखें↗

    यह प्रोजेक्ट TensorFlow पर आधारित एक वॉइस कन्वर्जन फ्रेमवर्क और डीप लर्निंग ऑडियो टूलकिट है, जिसे न्यूरल वॉइस स्टाइल ट्रांसफर के लिए बनाया गया है। यह एक स्पीच सिंथेसिस इंजन के रूप में काम करता है जो सोर्स स्पीकर की आवाज़ की स्पेक्ट्रल विशेषताओं को टारगेट स्पीकर की आवाज़ में बदल देता है। यह सिस्टम वॉइस कन्वर्जन के लिए फोनम-आधारित (phoneme-based) दृष्टिकोण अपनाता है, जो ऑडियो को स्पीकर-इंडिपेंडेंट फोनम में वर्गीकृत करता है और फिर उन्हें टारगेट वॉइस का उपयोग करके फिर से सिंथेसाइज करता है। यह पाइपलाइन अलग-अलग स्पीकर्स के बीच ऑडियो फीचर्स को मैप करके वॉइस विशेषताओं को बदलने की सुविधा देती है। इस टूलकिट में मल्टीपल GPUs पर ऑडियो मॉडल ट्रेनिंग, टेंसर डेटा नॉर्मलाइजेशन और मॉडल हाइपरपैरामीटर्स के प्रबंधन की क्षमताएं शामिल हैं। यह परफॉरमेंस मॉनिटरिंग के लिए भी टूल्स प्रदान करता है, जैसे कि कन्फ्यूजन मैट्रिक्स के जरिए क्लासिफिकेशन एक्यूरेसी को विज़ुअलाइज़ करना।

    Provides a toolkit for training voice models and normalizing tensor data using neural architectures.

    Python
    GitHub पर देखें↗3,941
  • riffusion/riffusion-hobbyriffusion का अवतार

    riffusion/riffusion-hobby

    3,895GitHub पर देखें↗

    Riffusion-hobby एक जनरेटिव AI टूल है जो Stable Diffusion के माध्यम से स्पेक्ट्रोग्राम इमेज बनाकर और उन्हें बजाने योग्य ऑडियो में बदलकर संगीत बनाता है। यह एक स्पेक्ट्रोग्राम ऑडियो सिंथेसाइज़र के रूप में कार्य करता है, जो ध्वनि के इमेज-आधारित आवृत्ति अभ्यासों को ऑडियो फ़ाइलों में बदलने के लिए डीप लर्निंग का उपयोग करता है। यह प्रोजेक्ट एक AI संगीत अनुमान सर्वर के रूप में संचालित होता है, जो टेक्स्ट प्रॉम्प्ट और सीड इमेज से ऑडियो उत्पन्न करने के लिए एक वेब-आधारित API एंडपॉइंट प्रदान करता है। इसमें संगीत निर्माण कार्यों को निष्पादित करने और स्वचालित ऑडियो निर्माण के लिए डिफ्यूजन मॉडल को कॉन्फ़िगर करने के लिए एक कमांड लाइन इंटरफ़ेस, साथ ही ध्वनि अभ्यासों में हेरफेर करने के लिए एक रीयल-टाइम ऑडियो जनरेटर भी शामिल है। सिस्टम क्लाउड मॉडल परिनियोजन, रिमोट अनुमान होस्टिंग और इमेज-टू-ऑडियो रूपांतरण के लिए डिजिटल सिग्नल प्रोसेसिंग सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह मॉडल पैरामीटर्स के साथ प्रयोग करने और संगीत निर्माण सेटिंग्स की खोज करने के लिए एक इंटरैक्टिव वेब-आधारित प्लेग्राउंड भी प्रदान करता है।

    Uses deep learning models to convert image-based spectrograms into playable audio files.

    Pythonaiaudiodiffusers
    GitHub पर देखें↗3,895
  • stability-ai/stable-audio-toolsStability-AI का अवतार

    Stability-AI/stable-audio-tools

    3,790GitHub पर देखें↗

    Stable-audio-tools is a toolkit for training and deploying latent diffusion models for high-fidelity audio synthesis. It provides a framework for generating audio by iteratively refining noise within a compressed latent space, using specialized encoders to preserve temporal and spectral features of the audio signal. The project features a system for adapting pre-trained audio checkpoints to new datasets through modular initialization and configuration files. It includes utilities for weight extraction and inference model export, which remove training metadata and optimizer states to create li

    Implements a deep learning library for high-fidelity audio synthesis and processing using neural architectures.

    Python
    GitHub पर देखें↗3,790
  • xiph/opusxiph का अवतार

    xiph/opus

    3,035GitHub पर देखें↗

    Opus is a lossy audio compression standard and codec designed for high-quality speech and music transmission over the internet. It functions as a low-latency audio codec and network-resilient streamer, providing a framework for encoding and decoding digital audio. The project distinguishes itself through the support of multi-channel ambisonics for immersive three-dimensional spatial audio reproduction. It is specifically optimized for real-time interactive communication, utilizing dynamic bitrate adjustment and forward error correction to maintain audio quality on unstable networks. The syst

    Embeds recovery data within packet padding using deep learning to maintain audio quality across lossy networks.

    Caudioccodec
    GitHub पर देखें↗3,035
  • pannous/tensorflow-speech-recognitionpannous का अवतार

    pannous/tensorflow-speech-recognition

    2,172GitHub पर देखें↗

    यह लाइब्रेरी स्पीच रिकग्निशन और ऑडियो क्लासिफिकेशन करने के लिए न्यूरल नेटवर्क को प्रशिक्षित करने के लिए एक डीप लर्निंग फ्रेमवर्क प्रदान करती है। यह वेरिएबल-लेंथ ऑडियो इनपुट को टेक्स्ट या संख्यात्मक आउटपुट में मैप करने के लिए सीक्वेंस-टू-सीक्वेंस आर्किटेक्चर का उपयोग करती है, जो कस्टम स्पीच-टू-टेक्स्ट ट्रांसक्रिप्शन मॉडल्स के विकास को सक्षम बनाती है। यह प्रोजेक्ट एकीकृत ऑडियो प्रोसेसिंग क्षमताओं के माध्यम से अलग है जो रॉ वेवफॉर्म्स को स्पेक्ट्रोग्राम और उच्च-आयामी संख्यात्मक वैक्टर में बदलती है। ये टूल्स वक्ताओं की पहचान करने के लिए अद्वितीय मुखर विशेषताओं के निष्कर्षण, साथ ही विशिष्ट ऑडियो सोर्सेज और बोले गए अंकों के वर्गीकरण की अनुमति देते हैं। मॉडल विकास का समर्थन करने के लिए, लाइब्रेरी में ऑडियो ऑगमेंटेशन और सिग्नल पुनर्निर्माण के लिए उपयोगिताएं शामिल हैं। विभिन्न ध्वनिक एनवायरनमेंट का अनुकरण करने के लिए ऑडियो नमूनों को प्रोग्रामेटिक रूप से संशोधित करके और लेटेंट-स्पेस पुनर्निर्माण के माध्यम से सीखे गए फीचर्स की अखंडता को सत्यापित करके, सिस्टम अपने अंतर्निहित न्यूरल नेटवर्क की मजबूती में सुधार करता है।

    Regenerates original spectrograms from compressed latent representations to verify the quality of learned audio features.

    Pythondeep-learningneural-networkspeech-recognition
    GitHub पर देखें↗2,172
  • voice-cloning-app/voice-cloning-appvoice-cloning-app का अवतार

    voice-cloning-app/Voice-Cloning-App

    1,438GitHub पर देखें↗

    This application is a platform for AI voice synthesis and neural voice cloning. It provides a comprehensive toolkit for converting text into natural-sounding human speech by applying custom-trained neural network models to specific audio samples. The system facilitates the entire lifecycle of voice model development, including the preparation of raw audiobooks and video transcriptions into structured training datasets. It supports the training of these models on local or remote hardware, utilizing multi-GPU distributed processing to handle large-scale data and accelerate model convergence. B

    Provides a toolkit for managing and processing large-scale voice datasets to facilitate high-performance speech synthesis.

    Pythondeep-learningpythonpytorch
    GitHub पर देखें↗1,438
  • bytedance/music_source_separationbytedance का अवतार

    bytedance/music_source_separation

    1,385GitHub पर देखें↗

    This project is a deep learning toolkit designed for audio source separation and music information retrieval. It provides a framework for decomposing polyphonic audio signals into distinct components, such as vocals, drums, and bass, by processing raw waveforms through neural network architectures. The library enables users to train custom separation models or fine-tune existing ones to improve accuracy on specific audio datasets. It supports the entire model lifecycle, including the conversion of raw audio into structured, indexed formats to optimize data loading and training efficiency. Th

    Provides a toolkit for training and fine-tuning neural networks to perform complex signal separation on audio waveforms.

    Pythonresearch
    GitHub पर देखें↗1,385
  1. Home
  2. Artificial Intelligence & ML
  3. Deep Learning Audio Libraries

सब-टैग एक्सप्लोर करें

  • Audio Synthesis Models1 सब-टैगNeural network models specifically designed to generate high-fidelity audio waveforms. **Distinct from Deep Learning Audio Libraries:** Focuses on the model implementation for audio generation rather than a general-purpose library of audio tools.
  • Neural Audio Reconstruction1 सब-टैगUsing deep learning to fill gaps or recover audio quality in lossy streams. **Distinct from Deep Learning Audio Libraries:** Distinct from Deep Learning Audio Libraries: specifically applied to the problem of packet loss recovery in network streams.
  • Voice Isolation ModelsNeural network models specifically designed to separate human speech from ambient background noise. **Distinct from Deep Learning Audio Libraries:** Focuses on the application of speech isolation rather than general audio synthesis or library infrastructure.