awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

14 个仓库

Awesome GitHub RepositoriesDeep Learning Audio Libraries

Libraries focused on high-fidelity audio synthesis and processing using neural architectures.

Distinct from Deep Learning Libraries: Shortlist targets general DL libraries or simple audio processing; this is a deep learning library for synthesis.

Explore 14 awesome GitHub repositories matching artificial intelligence & ml · Deep Learning Audio Libraries. Refine with filters or upvote what's useful.

Awesome Deep Learning Audio Libraries GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • deezer/spleeterdeezer 的头像

    deezer/spleeter

    28,252在 GitHub 上查看↗

    Spleeter is an AI audio source separation library and deep learning toolkit designed to split mixed music files into individual audio stems, such as vocals and drums. It provides a suite of pretrained models for isolating different instruments and voices from a recording. The toolkit includes capabilities for training and evaluating custom audio separation models using labeled datasets and configuration files. It also features utilities for measuring model performance by comparing separation outputs against reference datasets. The system manages audio processing through spectral representati

    Functions as a deep learning library for splitting music files into individual audio stems.

    Pythonaudio-processingbassdeep-learning
    在 GitHub 上查看↗28,252
  • facebookresearch/audiocraftfacebookresearch 的头像

    facebookresearch/audiocraft

    23,379在 GitHub 上查看↗

    Audiocraft is a deep learning audio library and machine learning framework designed for training, fine-tuning, and evaluating generative models for music and sound effects. It functions as a text-to-music generative model and a neural audio codec, providing the tools necessary to compress audio signals into discrete representations and synthesize high-fidelity waveforms from textual descriptions. The framework is distinguished by its ability to combine multiple conditioning signals, allowing for the generation of audio based on text prompts, melodic excerpts, or style-based audio clips. It al

    Ships a library for processing and generating high-fidelity audio using neural networks and transformer models.

    Jupyter Notebook
    在 GitHub 上查看↗23,379
  • wwmm/easyeffectswwmm 的头像

    wwmm/easyeffects

    9,690在 GitHub 上查看↗

    EasyEffects is a real-time audio processor and system-wide effects manager designed for PipeWire audio streams. It functions as a comprehensive suite for applying filters, equalizers, and limiters to both input and output audio across the entire system. The project distinguishes itself through its use of deep learning for neural network noise suppression and voice isolation, as well as its ability to simulate physical acoustic environments using impulse-response convolution. It includes a sophisticated preset management system that allows users to associate specific audio configurations with

    Uses deep learning models to isolate voice from ambient noise and preserve speech intelligibility.

    HTMLauto-volumecompressorequalizer
    在 GitHub 上查看↗9,690
  • jaywalnut310/vitsjaywalnut310 的头像

    jaywalnut310/vits

    7,862在 GitHub 上查看↗

    This project is an end-to-end text-to-speech engine and deep learning voice synthesizer. It functions as a neural speech synthesis framework that converts written text directly into audio waveforms using a single neural network. The system implements an adversarial framework and a conditional variational autoencoder to generate high-fidelity artificial speech. It utilizes a generative adversarial network to ensure synthesized audio is indistinguishable from real human speech. The toolkit provides capabilities for neural speech synthesis, text-to-audio generation, and the training of custom v

    Functions as a deep learning audio library for training high-fidelity speech models from text and audio.

    Pythondeep-learningpytorchspeech-synthesis
    在 GitHub 上查看↗7,862
  • ibab/tensorflow-wavenetibab 的头像

    ibab/tensorflow-wavenet

    5,432在 GitHub 上查看↗

    该项目是用于原始音频波形生成的 TensorFlow 实现。它作为一个条件语音合成模型,使用扩张卷积神经网络架构产生合成音频样本。 该系统通过在训练和生成过程中结合全局条件和分类标识符来支持自定义语音建模。这使得模型能够模仿特定说话人或为神经文本转语音应用生成独特的音频特征。 该框架涵盖了深度学习音频合成,包括音频数据集处理、从波形文件进行模型训练以及生成可播放的音频文件。它利用扩张因果卷积、mu-law 压扩和量化 Softmax 输出等技术组件来处理音频数据中的长距离依赖关系。

    Implements a deep learning library for high-fidelity audio synthesis using dilated convolutional architectures.

    Python
    在 GitHub 上查看↗5,432
  • microsoft/muzicmicrosoft 的头像

    microsoft/muzic

    4,928在 GitHub 上查看↗

    Muzic 是一个用于 AI 驱动的音乐分析、创作和合成的深度学习平台和框架。它作为一个音乐生成框架和分析工具,利用大型语言模型和自主智能体来编排符号音乐和音频音乐的创作与解读。 该项目以其跨模态能力而著称,将自然语言和符号音乐映射到共享的联合嵌入空间中,用于零样本分类和信息检索。它采用了多种专门的架构,包括用于音频合成的扩散框架、用于长序列结构一致性的双粒度注意力机制,以及结合音乐理论规则与神经网络的混合系统。 该平台涵盖了广泛的功能,包括从文本和歌词生成 MIDI 序列、神经歌声合成以及自动歌词转录。它还提供用于音乐结构建模、基于属性的符号生成以及通过自主智能体编排外部音乐工具的工具。 支持性实用程序包括用于大规模 MIDI 二进制化、数据集编码的数据工程流水线,以及用于旋律音符提取和语音到音素对齐的音频信号处理。

    Synthesizes musical audio and MIDI sequences using neural networks and deep learning models.

    Pythonai-musicdeep-learningmusic
    在 GitHub 上查看↗4,928
  • google/lyragoogle 的头像

    google/lyra

    3,964在 GitHub 上查看↗

    Lyra 是一个语音压缩框架和低比特率语音编解码器,旨在通过带宽受限的网络传输高质量音频。它利用自适应比特率音频编解码器在活动会话期间平衡音频质量和网络带宽。 该项目采用生成式音频压缩,使用神经网络从极少量数据中合成语音信号并重建缺失的音频细节。这允许从高度压缩的字节流中重建高质量语音音频。 该系统涵盖了带宽优化的 VoIP 和实时语音通信,专注于低比特率语音压缩以保持通话稳定性。其功能包括动态音频比特率调整和语音音频处理,以防止不稳定网络环境中的延迟和信号丢失。

    Recreates high-quality speech signals from compressed byte streams using deep learning models.

    C++
    在 GitHub 上查看↗3,964
  • andabi/deep-voice-conversionandabi 的头像

    andabi/deep-voice-conversion

    3,941在 GitHub 上查看↗

    这是一个基于 TensorFlow 的语音转换框架和深度学习音频工具包,专为神经语音风格迁移而设计。它作为一个语音合成引擎,将源说话人的语音频谱特征转换为目标说话人的声音特征。 该系统采用基于音素的语音转换方法,将音频话语分类为与说话人无关的音素,并使用目标声音重新合成它们。此流水线允许通过映射不同说话人之间的音频特征来转换语音特征。 该工具包包括跨多个 GPU 进行音频模型训练、张量数据归一化以及管理模型超参数的功能。它还提供了用于监控性能的工具,例如通过混淆矩阵可视化分类准确率。

    Provides a toolkit for training voice models and normalizing tensor data using neural architectures.

    Python
    在 GitHub 上查看↗3,941
  • riffusion/riffusion-hobbyriffusion 的头像

    riffusion/riffusion-hobby

    3,895在 GitHub 上查看↗

    Riffusion-hobby 是一款生成式 AI 工具,通过 Stable Diffusion 生成频谱图图像并将其转换为可播放的音频来创作音乐。它作为一个频谱图音频合成器,利用深度学习将基于图像的声音频率表示转换为音频文件。 该项目作为 AI 音乐推理服务器运行,提供基于 Web 的 API 端点,用于从文本提示词和种子图像生成音频。它还包括用于执行音乐生成任务、配置扩散模型以进行自动化音频创作的命令行界面,以及用于操作声音表示的实时音频生成器。 该系统涵盖了广泛的功能,包括云模型部署、远程推理托管以及用于图像转音频转换的数字信号处理。它还提供了一个交互式 Web 游乐场,用于试验模型参数和探索音乐生成设置。

    Uses deep learning models to convert image-based spectrograms into playable audio files.

    Pythonaiaudiodiffusers
    在 GitHub 上查看↗3,895
  • stability-ai/stable-audio-toolsStability-AI 的头像

    Stability-AI/stable-audio-tools

    3,790在 GitHub 上查看↗

    Stable-audio-tools is a toolkit for training and deploying latent diffusion models for high-fidelity audio synthesis. It provides a framework for generating audio by iteratively refining noise within a compressed latent space, using specialized encoders to preserve temporal and spectral features of the audio signal. The project features a system for adapting pre-trained audio checkpoints to new datasets through modular initialization and configuration files. It includes utilities for weight extraction and inference model export, which remove training metadata and optimizer states to create li

    Implements a deep learning library for high-fidelity audio synthesis and processing using neural architectures.

    Python
    在 GitHub 上查看↗3,790
  • xiph/opusxiph 的头像

    xiph/opus

    3,035在 GitHub 上查看↗

    Opus is a lossy audio compression standard and codec designed for high-quality speech and music transmission over the internet. It functions as a low-latency audio codec and network-resilient streamer, providing a framework for encoding and decoding digital audio. The project distinguishes itself through the support of multi-channel ambisonics for immersive three-dimensional spatial audio reproduction. It is specifically optimized for real-time interactive communication, utilizing dynamic bitrate adjustment and forward error correction to maintain audio quality on unstable networks. The syst

    Embeds recovery data within packet padding using deep learning to maintain audio quality across lossy networks.

    Caudioccodec
    在 GitHub 上查看↗3,035
  • pannous/tensorflow-speech-recognitionpannous 的头像

    pannous/tensorflow-speech-recognition

    2,172在 GitHub 上查看↗

    该库提供了一个深度学习框架,用于训练神经网络以执行语音识别和音频分类。它利用序列到序列(sequence-to-sequence)架构将变长音频输入映射为文本或数值输出,从而支持自定义语音转录模型的开发。 该项目通过集成的音频处理能力脱颖而出,这些能力将原始波形转换为频谱图和高维数值向量。这些工具允许提取独特的语音特征以识别说话人,以及对特定音频源和口述数字进行分类。 为支持模型开发,该库包含用于音频增强和信号重建的实用程序。通过以编程方式修改音频样本以模拟多样化的声学环境,并验证学习特征的完整性,该系统提高了其底层神经网络的鲁棒性。

    Regenerates original spectrograms from compressed latent representations to verify the quality of learned audio features.

    Pythondeep-learningneural-networkspeech-recognition
    在 GitHub 上查看↗2,172
  • voice-cloning-app/voice-cloning-appvoice-cloning-app 的头像

    voice-cloning-app/Voice-Cloning-App

    1,438在 GitHub 上查看↗

    This application is a platform for AI voice synthesis and neural voice cloning. It provides a comprehensive toolkit for converting text into natural-sounding human speech by applying custom-trained neural network models to specific audio samples. The system facilitates the entire lifecycle of voice model development, including the preparation of raw audiobooks and video transcriptions into structured training datasets. It supports the training of these models on local or remote hardware, utilizing multi-GPU distributed processing to handle large-scale data and accelerate model convergence. B

    Provides a toolkit for managing and processing large-scale voice datasets to facilitate high-performance speech synthesis.

    Pythondeep-learningpythonpytorch
    在 GitHub 上查看↗1,438
  • bytedance/music_source_separationbytedance 的头像

    bytedance/music_source_separation

    1,385在 GitHub 上查看↗

    This project is a deep learning toolkit designed for audio source separation and music information retrieval. It provides a framework for decomposing polyphonic audio signals into distinct components, such as vocals, drums, and bass, by processing raw waveforms through neural network architectures. The library enables users to train custom separation models or fine-tune existing ones to improve accuracy on specific audio datasets. It supports the entire model lifecycle, including the conversion of raw audio into structured, indexed formats to optimize data loading and training efficiency. Th

    Provides a toolkit for training and fine-tuning neural networks to perform complex signal separation on audio waveforms.

    Pythonresearch
    在 GitHub 上查看↗1,385
  1. Home
  2. Artificial Intelligence & ML
  3. Deep Learning Audio Libraries

探索子标签

  • Audio Synthesis Models1 个子标签Neural network models specifically designed to generate high-fidelity audio waveforms. **Distinct from Deep Learning Audio Libraries:** Focuses on the model implementation for audio generation rather than a general-purpose library of audio tools.
  • Neural Audio Reconstruction1 个子标签Using deep learning to fill gaps or recover audio quality in lossy streams. **Distinct from Deep Learning Audio Libraries:** Distinct from Deep Learning Audio Libraries: specifically applied to the problem of packet loss recovery in network streams.
  • Voice Isolation ModelsNeural network models specifically designed to separate human speech from ambient background noise. **Distinct from Deep Learning Audio Libraries:** Focuses on the application of speech isolation rather than general audio synthesis or library infrastructure.