awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

28 个仓库

Awesome GitHub RepositoriesAudio Generation Models

Models specialized in processing and generating high-quality audio content.

Distinguishing note: Focuses on audio-specific model architectures.

Explore 28 awesome GitHub repositories matching artificial intelligence & ml · Audio Generation Models. Refine with filters or upvote what's useful.

Awesome Audio Generation Models GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • facebookresearch/audiocraftfacebookresearch 的头像

    facebookresearch/audiocraft

    23,379在 GitHub 上查看↗

    Audiocraft is a deep learning audio library and machine learning framework designed for training, fine-tuning, and evaluating generative models for music and sound effects. It functions as a text-to-music generative model and a neural audio codec, providing the tools necessary to compress audio signals into discrete representations and synthesize high-fidelity waveforms from textual descriptions. The framework is distinguished by its ability to combine multiple conditioning signals, allowing for the generation of audio based on text prompts, melodic excerpts, or style-based audio clips. It al

    Generates subsequent audio content based on a provided starting audio clip to extend a sound sequence.

    Jupyter Notebook
    在 GitHub 上查看↗23,379
  • funaudiollm/cosyvoiceFunAudioLLM 的头像

    FunAudioLLM/CosyVoice

    21,673在 GitHub 上查看↗

    CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus

    Implements a neural synthesis architecture that modulates vocal attributes to produce speech with customizable emotional personas.

    Pythonaudio-generationcantonesechatbot
    在 GitHub 上查看↗21,673
  • nari-labs/dianari-labs 的头像

    nari-labs/dia

    19,324在 GitHub 上查看↗

    Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt

    Supports the injection of realistic nonverbal vocal expressions like laughter or gasps through specialized text markers.

    Pythonaiopen-weighttext-to-speech
    在 GitHub 上查看↗19,324
  • google-gemini/gemini-fullstack-langgraph-quickstartgoogle-gemini 的头像

    google-gemini/gemini-fullstack-langgraph-quickstart

    18,217在 GitHub 上查看↗

    This project is an agentic workflow orchestrator designed for building and deploying autonomous systems that perform multi-step reasoning. It functions as a tool-augmented engine, enabling developers to chain model calls with external function execution to complete complex, user-defined tasks. By integrating large language models with persistent memory and stateful logic, the framework supports the creation of intelligent applications capable of independent operation. The platform distinguishes itself through graph-based state orchestration, which allows developers to define logic steps and t

    Provides high-quality, low-latency audio-to-audio models for real-time interaction.

    Jupyter Notebookgeminigemini-api
    在 GitHub 上查看↗18,217
  • google-gemini/cookbookgoogle-gemini 的头像

    google-gemini/cookbook

    17,418在 GitHub 上查看↗

    The Gemini Cookbook is a comprehensive collection of implementation patterns, code samples, and development guides designed for building applications with Google Gemini models. It serves as a central resource for developers to integrate multimodal generative artificial intelligence into their software, providing the necessary frameworks to manage model interactions, stateful workflows, and structured data extraction. The repository distinguishes itself by offering specialized toolkits for autonomous agent orchestration, enabling the construction of agents that can execute code, browse the web

    Creates high-fidelity stereo music from text or image inputs, supporting custom lyrics and multi-language vocal performances.

    Jupyter Notebookgeminigemini-api
    在 GitHub 上查看↗17,418
  • sesameailabs/csmSesameAILabs 的头像

    SesameAILabs/csm

    14,669在 GitHub 上查看↗

    CSM is a conversational speech generation model and text-to-speech engine that converts text and audio inputs into synthetic speech. It utilizes a large language model architecture to predict and decode audio tokens for voice synthesis. The system functions as a zero-shot voice cloner, replicating specific speaker identities using short audio samples without requiring additional training. This enables precise control over speaker identity and the creation of synthetic speech that mimics a specific person. The model covers conversational speech synthesis and text-to-speech generation, transfo

    Utilizes a large language model architecture to predict and decode audio tokens for voice synthesis.

    Python
    在 GitHub 上查看↗14,669
  • alibaba/mnnalibaba 的头像

    alibaba/MNN

    14,242在 GitHub 上查看↗

    MNN is a high-performance inference engine and framework designed for on-device machine learning. It provides a comprehensive environment for executing, optimizing, and deploying neural network models directly on mobile and resource-constrained edge devices. The framework distinguishes itself through a robust model optimization toolkit that supports quantization, compression, and structural graph manipulation to minimize memory footprint and maximize execution speed. It features a modular architecture that abstracts hardware-specific backends, allowing models to run efficiently across diverse

    Generates audio waveform data from multimodal models for real-time streaming or file-based synthesis.

    C++armconvolutiondeep-learning
    在 GitHub 上查看↗14,242
  • winfredy/sadtalkerWinfredy 的头像

    Winfredy/SadTalker

    13,919在 GitHub 上查看↗

    SadTalker is a generative framework designed to synthesize expressive talking head videos from static portrait images. By mapping audio signals or text prompts to three-dimensional facial motion coefficients, the system synchronizes lip movements, facial expressions, and head orientation to create realistic digital character performances. The project distinguishes itself by decoupling identity from dynamic motion through latent space encoding, ensuring that the generated animations maintain visual fidelity to the source portrait. It supports comprehensive motion synthesis, including full-body

    Maps input audio signals to three-dimensional facial coefficients to synchronize lip movements and expressions with the source portrait.

    Python
    在 GitHub 上查看↗13,919
  • opentalker/sadtalkerOpenTalker 的头像

    OpenTalker/SadTalker

    13,895在 GitHub 上查看↗

    SadTalker is an audio-driven talking head generator that produces synchronized speaking videos from a single source image and an input audio file. The system utilizes a deep learning framework to map speech signals to facial motion data, enabling the creation of lifelike digital avatars and animated characters. The project distinguishes itself by employing a three-dimensional morphable model to translate audio features into precise facial landmarks and head pose parameters. It integrates latent diffusion motion synthesis to generate naturalistic head movements and uses expression-aware textur

    Extracts rhythmic and phonetic features from speech to drive the temporal evolution of facial expressions.

    Pythonaudio-driven-talking-facecvpr2023deep-fake
    在 GitHub 上查看↗13,895
  • facebookresearch/seamless_communicationfacebookresearch 的头像

    facebookresearch/seamless_communication

    11,797在 GitHub 上查看↗

    This project is a multimodal translation framework and large language model capable of speech-to-speech, speech-to-text, and text-to-text translation across nearly 100 languages. It provides a real-time speech translation engine and a comprehensive toolkit for converting spoken audio between languages. The system is distinguished by its ability to preserve the original speaker's tone, pace, and prosody during translation. It utilizes a specialized on-device inference toolkit that converts model checkpoints into C-based libraries, enabling low-latency execution on mobile and edge hardware with

    Discovers pairs of audio segments across languages that share both the same meaning and expressivity.

    Jupyter Notebook
    在 GitHub 上查看↗11,797
  • aigc-audio/audiogptAIGC-Audio 的头像

    AIGC-Audio/AudioGPT

    10,174在 GitHub 上查看↗

    AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre

    Implements a system that generates soundscapes and audio clips from visual images or natural language descriptions.

    Pythonaudiogptmusic
    在 GitHub 上查看↗10,174
  • open-mmlab/amphionopen-mmlab 的头像

    open-mmlab/Amphion

    9,844在 GitHub 上查看↗

    Amphion is an audio generation toolkit designed for the research and development of models that synthesize speech, music, and environmental sound effects. It provides a standardized framework for reproducible audio synthesis, incorporating a text-to-speech engine and a voice conversion framework. The project specializes in transforming audio identities, allowing for the modification of speaker accents and voice identities while preserving original rhythm and style. It also includes capabilities for singing voice synthesis and the generation of environmental soundscapes from text descriptions

    Produces high-quality audio waveforms from intermediate representations using specialized neural vocoders.

    Pythonaudio-generationaudio-synthesisaudioldm
    在 GitHub 上查看↗9,844
  • ml-explore/mlx-examplesml-explore 的头像

    ml-explore/mlx-examples

    8,254在 GitHub 上查看↗

    This repository provides a collection of reference implementations and code examples for training and deploying machine learning models using the MLX framework. It serves as a practical guide for executing distributed training, fine-tuning large language models, converting model weights, and implementing multimodal generative workflows. The project distinguishes itself through specialized examples for local hardware execution, featuring weight quantization to reduce memory usage and low-rank adaptation for parameter-efficient fine-tuning. It also includes scripts for transforming external mod

    Implements generative workflows for producing text, images, audio, and video from mixed-modal inputs.

    Pythonmlx
    在 GitHub 上查看↗8,254
  • openai/jukeboxopenai 的头像

    openai/jukebox

    8,039在 GitHub 上查看↗

    Jukebox is a generative audio model and AI music synthesis tool designed to create high-fidelity music samples and singing voices. It functions as a deep learning system that synthesizes raw audio conditioned on genre and artist metadata, utilizing a neural audio codec to convert raw audio into discrete codes for generative modeling and reconstruction. The system enables musical style steering and AI music composition by conditioning generation on specific artists, genres, and lyrics. It supports audio priming, allowing existing wave files to guide the creation of new musical sequences, and p

    Implements a generative audio model that predicts sequences of audio codes to create music.

    Pythonaudiogenerative-modelmusic
    在 GitHub 上查看↗8,039
  • humanaigc/emoHumanAIGC 的头像

    HumanAIGC/EMO

    7,616在 GitHub 上查看↗

    EMO 是一个 AI 人像动画和音频转视频扩散模型,旨在生成富有表现力的说话人视频。它将单张静态人像图片和一段音频轨道转换为同步的说话人视频。 该系统专注于数字人合成,产生高保真的面部动作和情感线索。它将唇部动作和面部表情与语音录音同步,从而创建逼真的人像动画。 该框架利用扩散过程和跨模态对齐机制来确保音频信号与视觉关键点之间的时间同步。它采用基于参考的图像调节来保持身份一致性,并使用时间一致性层来确保帧间运动的流畅性。

    Generates high-fidelity human facial movements and emotional cues driven by audio signals.

    在 GitHub 上查看↗7,616
  • ibab/tensorflow-wavenetibab 的头像

    ibab/tensorflow-wavenet

    5,432在 GitHub 上查看↗

    该项目是用于原始音频波形生成的 TensorFlow 实现。它作为一个条件语音合成模型,使用扩张卷积神经网络架构产生合成音频样本。 该系统通过在训练和生成过程中结合全局条件和分类标识符来支持自定义语音建模。这使得模型能够模仿特定说话人或为神经文本转语音应用生成独特的音频特征。 该框架涵盖了深度学习音频合成,包括音频数据集处理、从波形文件进行模型训练以及生成可播放的音频文件。它利用扩张因果卷积、mu-law 压扩和量化 Softmax 输出等技术组件来处理音频数据中的长距离依赖关系。

    Provides a model capable of producing high-quality raw audio waveforms from trained weights.

    Python
    在 GitHub 上查看↗5,432
  • zejun-yang/aniportraitZejun-Yang 的头像

    Zejun-Yang/AniPortrait

    5,020在 GitHub 上查看↗

    AniPortrait 是一个 AI 视频合成流水线,旨在生成照片级逼真的说话肖像和面部动画。它充当说话头像生成器和音频驱动的动画师,将唇部动作、表情和头部姿势与语音或参考视频源同步。 该系统包括一个面部表情迁移工具,用于将源视频中的动作重演到静态参考图像上。它利用带有参考图像调节的潜在扩散模型,在生成的帧中保持视觉身份和一致性。 该流水线涵盖音频到表情的映射、姿势引导的运动控制和照片级逼真的视频合成。它结合了帧插值上采样,以加速生成过程并减少总渲染时间。

    Implements neural encoders that extract features from audio to drive facial expression parameters.

    Python
    在 GitHub 上查看↗5,020
  • moonintheriver/diffsingerMoonInTheRiver 的头像

    MoonInTheRiver/DiffSinger

    4,804在 GitHub 上查看↗

    DiffSinger is an AI vocal synthesizer and neural audio generator designed to produce high-fidelity singing and speech. It functions as a text-to-speech system and a diffusion-based singing voice synthesis tool that transforms text and pitch into audible audio. The system utilizes a shallow diffusion mechanism and iterative noise refinement to generate realistic vocal performances. It incorporates specialized sampling plugins and numerical solvers to accelerate inference and reduce the time required to generate synthetic voices. The project covers acoustic modeling, mel-spectrogram synthesis,

    Produces high-fidelity audio waveforms and spectrograms for both singing and spoken language.

    Pythonaaai2022diffusion-modeldiffusion-speedup
    在 GitHub 上查看↗4,804
  • huggingface/diffusion-models-classhuggingface 的头像

    huggingface/diffusion-models-class

    4,331在 GitHub 上查看↗

    本项目是一个专注于生成式扩散模型的教育课程和培训材料集合。它提供了一个课程大纲和实践指南,用于训练、微调和部署能够合成图像、音频和视频的模型。 该材料涵盖了特定的实现策略,包括基于噪声的合成、迭代细化和潜在空间压缩。它提供了关于通过条件合成和提示词遵循优化来引导生成式输出的指导,以及图像修复和基于文本的编辑技术。 该项目包括关于模型优化和开发的内容,涵盖概念微调和推理步骤的减少。它还提供了用于生成合成媒体的工作流,例如生成视频序列和将视觉频谱图转换为音频。 实践实现通过 PyTorch 代码示例和将模型权重及元数据发布到 Hugging Face Hub 的教程提供。

    Provides a workflow for producing audio by generating visual spectrograms and converting them back into sound.

    Jupyter Notebook
    在 GitHub 上查看↗4,331
  • badtobest/echomimicBadToBest 的头像

    BadToBest/EchoMimic

    4,258在 GitHub 上查看↗

    EchoMimic 是一个音频驱动的肖像动画框架和潜在扩散视频生成器。它通过将面部动作与音轨和运动驱动程序同步,将静态参考图像转换为动态的说话头像视频。 该系统作为一个混合运动合成引擎,结合了音频输入和姿势数据。它利用面部标志运动控制器来编辑定位标记,从而实现精确的同步和视频到视频的姿势迁移。 该管道通过潜在扩散和面部标志调节涵盖了图像到视频的动画。这允许由音频、姿势或两种引导源的组合驱动的肖像动画。

    Implements neural encoders that translate speech features into facial expression parameters for synchronization.

    Python
    在 GitHub 上查看↗4,258
上一个12下一个
  1. Home
  2. Artificial Intelligence & ML
  3. Audio Generation Models

探索子标签

  • Audio Prompt ContinuationCapabilities for extending an existing audio sequence based on a starting clip. **Distinct from Audio Generation Models:** Focuses on sequence continuation specifically rather than general generation
  • Audio Sample Reconstruction1 个子标签Processes that convert model samples into final audio waveforms with configurable formats. **Distinct from Audio Generation Models:** Focuses on the reconstruction of waveforms from model samples
  • Expressive Synthesis Models3 个子标签Neural architectures that modulate vocal attributes and emotional personas during speech generation. **Distinct from Audio Generation Models:** Distinct from general audio generation: focuses specifically on expressive speech synthesis with emotional persona modulation.
  • Image-to-Audio Synthesis2 个子标签Specific neural models that translate visual content and context from images into corresponding audio. **Distinct from Audio Generation Models:** Narrowly focused on image-based triggers for audio generation, distinct from general audio generation models.
  • Integrity ValidationsVerification processes to ensure model outputs remain consistent with reference implementations. **Distinct from Audio Generation Models:** Focuses on correctness verification of processed waveforms rather than the generation process itself.
  • Melodic ConditioningGeneration of musical content guided by an existing audio melody. **Distinct from Audio Generation Models:** Specific to using melody as a guidance signal for generation
  • Multimodal GenerationAudio generation models that utilize non-audio inputs such as images or text to synthesize soundscapes. **Distinct from Audio Generation Models:** Focuses on cross-modal input mapping specifically, whereas Audio Generation Models is the broader category for any audio synthesis.