14 रिपॉजिटरी
Libraries focused on high-fidelity audio synthesis and processing using neural architectures.
Distinct from Deep Learning Libraries: Shortlist targets general DL libraries or simple audio processing; this is a deep learning library for synthesis.
Explore 14 awesome GitHub repositories matching artificial intelligence & ml · Deep Learning Audio Libraries. Refine with filters or upvote what's useful.
Spleeter is an AI audio source separation library and deep learning toolkit designed to split mixed music files into individual audio stems, such as vocals and drums. It provides a suite of pretrained models for isolating different instruments and voices from a recording. The toolkit includes capabilities for training and evaluating custom audio separation models using labeled datasets and configuration files. It also features utilities for measuring model performance by comparing separation outputs against reference datasets. The system manages audio processing through spectral representati
Functions as a deep learning library for splitting music files into individual audio stems.
Audiocraft is a deep learning audio library and machine learning framework designed for training, fine-tuning, and evaluating generative models for music and sound effects. It functions as a text-to-music generative model and a neural audio codec, providing the tools necessary to compress audio signals into discrete representations and synthesize high-fidelity waveforms from textual descriptions. The framework is distinguished by its ability to combine multiple conditioning signals, allowing for the generation of audio based on text prompts, melodic excerpts, or style-based audio clips. It al
Ships a library for processing and generating high-fidelity audio using neural networks and transformer models.
EasyEffects is a real-time audio processor and system-wide effects manager designed for PipeWire audio streams. It functions as a comprehensive suite for applying filters, equalizers, and limiters to both input and output audio across the entire system. The project distinguishes itself through its use of deep learning for neural network noise suppression and voice isolation, as well as its ability to simulate physical acoustic environments using impulse-response convolution. It includes a sophisticated preset management system that allows users to associate specific audio configurations with
Uses deep learning models to isolate voice from ambient noise and preserve speech intelligibility.
This project is an end-to-end text-to-speech engine and deep learning voice synthesizer. It functions as a neural speech synthesis framework that converts written text directly into audio waveforms using a single neural network. The system implements an adversarial framework and a conditional variational autoencoder to generate high-fidelity artificial speech. It utilizes a generative adversarial network to ensure synthesized audio is indistinguishable from real human speech. The toolkit provides capabilities for neural speech synthesis, text-to-audio generation, and the training of custom v
Functions as a deep learning audio library for training high-fidelity speech models from text and audio.
यह प्रोजेक्ट रॉ ऑडियो वेवफॉर्म जनरेशन के लिए एक न्यूरल नेटवर्क का TensorFlow कार्यान्वयन है। यह एक कंडीशन-आधारित स्पीच सिंथेसिस मॉडल के रूप में कार्य करता है जो डायलेटेड कन्वेन्शनल न्यूरल नेटवर्क आर्किटेक्चर का उपयोग करके सिंथेटिक ऑडियो सैंपल तैयार करता है। यह सिस्टम ट्रेनिंग और जनरेशन के दौरान ग्लोबल कंडीशनिंग और कैटेगोरिकल आइडेंटिफायर्स को शामिल करके कस्टम वॉयस मॉडलिंग का समर्थन करता है। यह मॉडल को न्यूरल टेक्स्ट-टू-स्पीच एप्लिकेशन के लिए विशिष्ट वक्ताओं या अलग-अलग ऑडियो विशेषताओं की नकल करने की अनुमति देता है। यह फ्रेमवर्क डीप लर्निंग ऑडियो सिंथेसिस को कवर करता है, जिसमें ऑडियो डेटासेट प्रोसेसिंग, वेवफॉर्म फाइलों से मॉडल ट्रेनिंग, और चलाने योग्य ऑडियो फाइलों का जनरेशन शामिल है। यह ऑडियो डेटा में लॉन्ग-रेंज डिपेंडेंसीज को संभालने के लिए डायलेटेड कॉज़ल कन्वेन्शन्स, mu-law कंपैंडिंग और क्वांटाइज्ड सॉफ्टमैक्स आउटपुट जैसे तकनीकी घटकों का उपयोग करता है।
Implements a deep learning library for high-fidelity audio synthesis using dilated convolutional architectures.
Muzic AI-संचालित संगीत विश्लेषण, रचना और संश्लेषण के लिए एक डीप लर्निंग प्लेटफ़ॉर्म और फ्रेमवर्क है। यह एक संगीत जनरेशन फ्रेमवर्क और विश्लेषण टूल के रूप में कार्य करता है, जो प्रतीकात्मक और ऑडियो संगीत के निर्माण और व्याख्या को व्यवस्थित करने के लिए बड़े भाषा मॉडल्स और स्वायत्त एजेंटों का उपयोग करता है। यह प्रोजेक्ट अपनी क्रॉस-मॉडल क्षमताओं द्वारा प्रतिष्ठित है, जो ज़ीरो-शॉट वर्गीकरण और सूचना पुनर्प्राप्ति के लिए प्राकृतिक भाषा और प्रतीकात्मक संगीत को एक साझा संयुक्त एम्बेडिंग स्पेस में मैप करता है। यह विभिन्न प्रकार के विशेष आर्किटेक्चर को नियोजित करता है, जिसमें ऑडियो संश्लेषण के लिए डिफ्यूज़न फ्रेमवर्क, लंबी-अनुक्रम संरचनात्मक स्थिरता के लिए डुअल-ग्रेन अटेंशन मैकेनिज्म और एक हाइब्रिड सिस्टम शामिल है जो न्यूरल नेटवर्क के साथ संगीत सिद्धांत नियमों को जोड़ता है। यह प्लेटफ़ॉर्म टेक्स्ट और लिरिक्स से MIDI अनुक्रमों के निर्माण, न्यूरल सिंगिंग वॉयस सिंथेसिस और स्वचालित लिरिक्स ट्रांसक्रिप्शन सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह संगीत संरचना मॉडलिंग, विशेषता-आधारित प्रतीकात्मक जनरेशन और स्वायत्त एजेंटों के माध्यम से बाहरी संगीत टूल्स के ऑर्केस्ट्रेशन के लिए टूल्स भी प्रदान करता है। सहायक यूटिलिटीज में बड़े पैमाने पर MIDI बाइनराइजेशन, डेटासेट एन्कोडिंग और मेलोडी नोट निष्कर्षण और स्पीच-टू-फोनम एलाइनमेंट के लिए ऑडियो सिग्नल प्रोसेसिंग के लिए डेटा इंजीनियरिंग पाइपलाइन शामिल हैं।
Synthesizes musical audio and MIDI sequences using neural networks and deep learning models.
Lyra is a voice compression framework and low-bitrate speech codec designed to transmit high-quality audio over bandwidth-constrained networks. It utilizes an adaptive bitrate audio codec to balance audio quality and network bandwidth during active sessions. The project employs generative audio compression, using neural networks to synthesize speech signals from minimal data and reconstruct missing audio details. This allows for high-quality voice audio reconstruction from highly compressed byte streams. The system covers bandwidth-optimized voice over IP and real-time voice communication, f
Recreates high-quality speech signals from compressed byte streams using deep learning models.
यह प्रोजेक्ट TensorFlow पर आधारित एक वॉइस कन्वर्जन फ्रेमवर्क और डीप लर्निंग ऑडियो टूलकिट है, जिसे न्यूरल वॉइस स्टाइल ट्रांसफर के लिए बनाया गया है। यह एक स्पीच सिंथेसिस इंजन के रूप में काम करता है जो सोर्स स्पीकर की आवाज़ की स्पेक्ट्रल विशेषताओं को टारगेट स्पीकर की आवाज़ में बदल देता है। यह सिस्टम वॉइस कन्वर्जन के लिए फोनम-आधारित (phoneme-based) दृष्टिकोण अपनाता है, जो ऑडियो को स्पीकर-इंडिपेंडेंट फोनम में वर्गीकृत करता है और फिर उन्हें टारगेट वॉइस का उपयोग करके फिर से सिंथेसाइज करता है। यह पाइपलाइन अलग-अलग स्पीकर्स के बीच ऑडियो फीचर्स को मैप करके वॉइस विशेषताओं को बदलने की सुविधा देती है। इस टूलकिट में मल्टीपल GPUs पर ऑडियो मॉडल ट्रेनिंग, टेंसर डेटा नॉर्मलाइजेशन और मॉडल हाइपरपैरामीटर्स के प्रबंधन की क्षमताएं शामिल हैं। यह परफॉरमेंस मॉनिटरिंग के लिए भी टूल्स प्रदान करता है, जैसे कि कन्फ्यूजन मैट्रिक्स के जरिए क्लासिफिकेशन एक्यूरेसी को विज़ुअलाइज़ करना।
Provides a toolkit for training voice models and normalizing tensor data using neural architectures.
Riffusion-hobby एक जनरेटिव AI टूल है जो Stable Diffusion के माध्यम से स्पेक्ट्रोग्राम इमेज बनाकर और उन्हें बजाने योग्य ऑडियो में बदलकर संगीत बनाता है। यह एक स्पेक्ट्रोग्राम ऑडियो सिंथेसाइज़र के रूप में कार्य करता है, जो ध्वनि के इमेज-आधारित आवृत्ति अभ्यासों को ऑडियो फ़ाइलों में बदलने के लिए डीप लर्निंग का उपयोग करता है। यह प्रोजेक्ट एक AI संगीत अनुमान सर्वर के रूप में संचालित होता है, जो टेक्स्ट प्रॉम्प्ट और सीड इमेज से ऑडियो उत्पन्न करने के लिए एक वेब-आधारित API एंडपॉइंट प्रदान करता है। इसमें संगीत निर्माण कार्यों को निष्पादित करने और स्वचालित ऑडियो निर्माण के लिए डिफ्यूजन मॉडल को कॉन्फ़िगर करने के लिए एक कमांड लाइन इंटरफ़ेस, साथ ही ध्वनि अभ्यासों में हेरफेर करने के लिए एक रीयल-टाइम ऑडियो जनरेटर भी शामिल है। सिस्टम क्लाउड मॉडल परिनियोजन, रिमोट अनुमान होस्टिंग और इमेज-टू-ऑडियो रूपांतरण के लिए डिजिटल सिग्नल प्रोसेसिंग सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। यह मॉडल पैरामीटर्स के साथ प्रयोग करने और संगीत निर्माण सेटिंग्स की खोज करने के लिए एक इंटरैक्टिव वेब-आधारित प्लेग्राउंड भी प्रदान करता है।
Uses deep learning models to convert image-based spectrograms into playable audio files.
Stable-audio-tools is a toolkit for training and deploying latent diffusion models for high-fidelity audio synthesis. It provides a framework for generating audio by iteratively refining noise within a compressed latent space, using specialized encoders to preserve temporal and spectral features of the audio signal. The project features a system for adapting pre-trained audio checkpoints to new datasets through modular initialization and configuration files. It includes utilities for weight extraction and inference model export, which remove training metadata and optimizer states to create li
Implements a deep learning library for high-fidelity audio synthesis and processing using neural architectures.
Opus is a lossy audio compression standard and codec designed for high-quality speech and music transmission over the internet. It functions as a low-latency audio codec and network-resilient streamer, providing a framework for encoding and decoding digital audio. The project distinguishes itself through the support of multi-channel ambisonics for immersive three-dimensional spatial audio reproduction. It is specifically optimized for real-time interactive communication, utilizing dynamic bitrate adjustment and forward error correction to maintain audio quality on unstable networks. The syst
Embeds recovery data within packet padding using deep learning to maintain audio quality across lossy networks.
यह लाइब्रेरी स्पीच रिकग्निशन और ऑडियो क्लासिफिकेशन करने के लिए न्यूरल नेटवर्क को प्रशिक्षित करने के लिए एक डीप लर्निंग फ्रेमवर्क प्रदान करती है। यह वेरिएबल-लेंथ ऑडियो इनपुट को टेक्स्ट या संख्यात्मक आउटपुट में मैप करने के लिए सीक्वेंस-टू-सीक्वेंस आर्किटेक्चर का उपयोग करती है, जो कस्टम स्पीच-टू-टेक्स्ट ट्रांसक्रिप्शन मॉडल्स के विकास को सक्षम बनाती है। यह प्रोजेक्ट एकीकृत ऑडियो प्रोसेसिंग क्षमताओं के माध्यम से अलग है जो रॉ वेवफॉर्म्स को स्पेक्ट्रोग्राम और उच्च-आयामी संख्यात्मक वैक्टर में बदलती है। ये टूल्स वक्ताओं की पहचान करने के लिए अद्वितीय मुखर विशेषताओं के निष्कर्षण, साथ ही विशिष्ट ऑडियो सोर्सेज और बोले गए अंकों के वर्गीकरण की अनुमति देते हैं। मॉडल विकास का समर्थन करने के लिए, लाइब्रेरी में ऑडियो ऑगमेंटेशन और सिग्नल पुनर्निर्माण के लिए उपयोगिताएं शामिल हैं। विभिन्न ध्वनिक एनवायरनमेंट का अनुकरण करने के लिए ऑडियो नमूनों को प्रोग्रामेटिक रूप से संशोधित करके और लेटेंट-स्पेस पुनर्निर्माण के माध्यम से सीखे गए फीचर्स की अखंडता को सत्यापित करके, सिस्टम अपने अंतर्निहित न्यूरल नेटवर्क की मजबूती में सुधार करता है।
Regenerates original spectrograms from compressed latent representations to verify the quality of learned audio features.
This application is a platform for AI voice synthesis and neural voice cloning. It provides a comprehensive toolkit for converting text into natural-sounding human speech by applying custom-trained neural network models to specific audio samples. The system facilitates the entire lifecycle of voice model development, including the preparation of raw audiobooks and video transcriptions into structured training datasets. It supports the training of these models on local or remote hardware, utilizing multi-GPU distributed processing to handle large-scale data and accelerate model convergence. B
Provides a toolkit for managing and processing large-scale voice datasets to facilitate high-performance speech synthesis.
This project is a deep learning toolkit designed for audio source separation and music information retrieval. It provides a framework for decomposing polyphonic audio signals into distinct components, such as vocals, drums, and bass, by processing raw waveforms through neural network architectures. The library enables users to train custom separation models or fine-tune existing ones to improve accuracy on specific audio datasets. It supports the entire model lifecycle, including the conversion of raw audio into structured, indexed formats to optimize data loading and training efficiency. Th
Provides a toolkit for training and fine-tuning neural networks to perform complex signal separation on audio waveforms.