28 个仓库
Models specialized in processing and generating high-quality audio content.
Distinguishing note: Focuses on audio-specific model architectures.
Explore 28 awesome GitHub repositories matching artificial intelligence & ml · Audio Generation Models. Refine with filters or upvote what's useful.
Audiocraft is a deep learning audio library and machine learning framework designed for training, fine-tuning, and evaluating generative models for music and sound effects. It functions as a text-to-music generative model and a neural audio codec, providing the tools necessary to compress audio signals into discrete representations and synthesize high-fidelity waveforms from textual descriptions. The framework is distinguished by its ability to combine multiple conditioning signals, allowing for the generation of audio based on text prompts, melodic excerpts, or style-based audio clips. It al
Generates subsequent audio content based on a provided starting audio clip to extend a sound sequence.
CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus
Implements a neural synthesis architecture that modulates vocal attributes to produce speech with customizable emotional personas.
Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt
Supports the injection of realistic nonverbal vocal expressions like laughter or gasps through specialized text markers.
This project is an agentic workflow orchestrator designed for building and deploying autonomous systems that perform multi-step reasoning. It functions as a tool-augmented engine, enabling developers to chain model calls with external function execution to complete complex, user-defined tasks. By integrating large language models with persistent memory and stateful logic, the framework supports the creation of intelligent applications capable of independent operation. The platform distinguishes itself through graph-based state orchestration, which allows developers to define logic steps and t
Provides high-quality, low-latency audio-to-audio models for real-time interaction.
The Gemini Cookbook is a comprehensive collection of implementation patterns, code samples, and development guides designed for building applications with Google Gemini models. It serves as a central resource for developers to integrate multimodal generative artificial intelligence into their software, providing the necessary frameworks to manage model interactions, stateful workflows, and structured data extraction. The repository distinguishes itself by offering specialized toolkits for autonomous agent orchestration, enabling the construction of agents that can execute code, browse the web
Creates high-fidelity stereo music from text or image inputs, supporting custom lyrics and multi-language vocal performances.
CSM is a conversational speech generation model and text-to-speech engine that converts text and audio inputs into synthetic speech. It utilizes a large language model architecture to predict and decode audio tokens for voice synthesis. The system functions as a zero-shot voice cloner, replicating specific speaker identities using short audio samples without requiring additional training. This enables precise control over speaker identity and the creation of synthetic speech that mimics a specific person. The model covers conversational speech synthesis and text-to-speech generation, transfo
Utilizes a large language model architecture to predict and decode audio tokens for voice synthesis.
MNN is a high-performance inference engine and framework designed for on-device machine learning. It provides a comprehensive environment for executing, optimizing, and deploying neural network models directly on mobile and resource-constrained edge devices. The framework distinguishes itself through a robust model optimization toolkit that supports quantization, compression, and structural graph manipulation to minimize memory footprint and maximize execution speed. It features a modular architecture that abstracts hardware-specific backends, allowing models to run efficiently across diverse
Generates audio waveform data from multimodal models for real-time streaming or file-based synthesis.
SadTalker is a generative framework designed to synthesize expressive talking head videos from static portrait images. By mapping audio signals or text prompts to three-dimensional facial motion coefficients, the system synchronizes lip movements, facial expressions, and head orientation to create realistic digital character performances. The project distinguishes itself by decoupling identity from dynamic motion through latent space encoding, ensuring that the generated animations maintain visual fidelity to the source portrait. It supports comprehensive motion synthesis, including full-body
Maps input audio signals to three-dimensional facial coefficients to synchronize lip movements and expressions with the source portrait.
SadTalker is an audio-driven talking head generator that produces synchronized speaking videos from a single source image and an input audio file. The system utilizes a deep learning framework to map speech signals to facial motion data, enabling the creation of lifelike digital avatars and animated characters. The project distinguishes itself by employing a three-dimensional morphable model to translate audio features into precise facial landmarks and head pose parameters. It integrates latent diffusion motion synthesis to generate naturalistic head movements and uses expression-aware textur
Extracts rhythmic and phonetic features from speech to drive the temporal evolution of facial expressions.
This project is a multimodal translation framework and large language model capable of speech-to-speech, speech-to-text, and text-to-text translation across nearly 100 languages. It provides a real-time speech translation engine and a comprehensive toolkit for converting spoken audio between languages. The system is distinguished by its ability to preserve the original speaker's tone, pace, and prosody during translation. It utilizes a specialized on-device inference toolkit that converts model checkpoints into C-based libraries, enabling low-latency execution on mobile and edge hardware with
Discovers pairs of audio segments across languages that share both the same meaning and expressivity.
AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre
Implements a system that generates soundscapes and audio clips from visual images or natural language descriptions.
Amphion is an audio generation toolkit designed for the research and development of models that synthesize speech, music, and environmental sound effects. It provides a standardized framework for reproducible audio synthesis, incorporating a text-to-speech engine and a voice conversion framework. The project specializes in transforming audio identities, allowing for the modification of speaker accents and voice identities while preserving original rhythm and style. It also includes capabilities for singing voice synthesis and the generation of environmental soundscapes from text descriptions
Produces high-quality audio waveforms from intermediate representations using specialized neural vocoders.
This repository provides a collection of reference implementations and code examples for training and deploying machine learning models using the MLX framework. It serves as a practical guide for executing distributed training, fine-tuning large language models, converting model weights, and implementing multimodal generative workflows. The project distinguishes itself through specialized examples for local hardware execution, featuring weight quantization to reduce memory usage and low-rank adaptation for parameter-efficient fine-tuning. It also includes scripts for transforming external mod
Implements generative workflows for producing text, images, audio, and video from mixed-modal inputs.
Jukebox is a generative audio model and AI music synthesis tool designed to create high-fidelity music samples and singing voices. It functions as a deep learning system that synthesizes raw audio conditioned on genre and artist metadata, utilizing a neural audio codec to convert raw audio into discrete codes for generative modeling and reconstruction. The system enables musical style steering and AI music composition by conditioning generation on specific artists, genres, and lyrics. It supports audio priming, allowing existing wave files to guide the creation of new musical sequences, and p
Implements a generative audio model that predicts sequences of audio codes to create music.
EMO 是一个 AI 人像动画和音频转视频扩散模型,旨在生成富有表现力的说话人视频。它将单张静态人像图片和一段音频轨道转换为同步的说话人视频。 该系统专注于数字人合成,产生高保真的面部动作和情感线索。它将唇部动作和面部表情与语音录音同步,从而创建逼真的人像动画。 该框架利用扩散过程和跨模态对齐机制来确保音频信号与视觉关键点之间的时间同步。它采用基于参考的图像调节来保持身份一致性,并使用时间一致性层来确保帧间运动的流畅性。
Generates high-fidelity human facial movements and emotional cues driven by audio signals.
该项目是用于原始音频波形生成的 TensorFlow 实现。它作为一个条件语音合成模型,使用扩张卷积神经网络架构产生合成音频样本。 该系统通过在训练和生成过程中结合全局条件和分类标识符来支持自定义语音建模。这使得模型能够模仿特定说话人或为神经文本转语音应用生成独特的音频特征。 该框架涵盖了深度学习音频合成,包括音频数据集处理、从波形文件进行模型训练以及生成可播放的音频文件。它利用扩张因果卷积、mu-law 压扩和量化 Softmax 输出等技术组件来处理音频数据中的长距离依赖关系。
Provides a model capable of producing high-quality raw audio waveforms from trained weights.
AniPortrait 是一个 AI 视频合成流水线,旨在生成照片级逼真的说话肖像和面部动画。它充当说话头像生成器和音频驱动的动画师,将唇部动作、表情和头部姿势与语音或参考视频源同步。 该系统包括一个面部表情迁移工具,用于将源视频中的动作重演到静态参考图像上。它利用带有参考图像调节的潜在扩散模型,在生成的帧中保持视觉身份和一致性。 该流水线涵盖音频到表情的映射、姿势引导的运动控制和照片级逼真的视频合成。它结合了帧插值上采样,以加速生成过程并减少总渲染时间。
Implements neural encoders that extract features from audio to drive facial expression parameters.
DiffSinger is an AI vocal synthesizer and neural audio generator designed to produce high-fidelity singing and speech. It functions as a text-to-speech system and a diffusion-based singing voice synthesis tool that transforms text and pitch into audible audio. The system utilizes a shallow diffusion mechanism and iterative noise refinement to generate realistic vocal performances. It incorporates specialized sampling plugins and numerical solvers to accelerate inference and reduce the time required to generate synthetic voices. The project covers acoustic modeling, mel-spectrogram synthesis,
Produces high-fidelity audio waveforms and spectrograms for both singing and spoken language.
本项目是一个专注于生成式扩散模型的教育课程和培训材料集合。它提供了一个课程大纲和实践指南,用于训练、微调和部署能够合成图像、音频和视频的模型。 该材料涵盖了特定的实现策略,包括基于噪声的合成、迭代细化和潜在空间压缩。它提供了关于通过条件合成和提示词遵循优化来引导生成式输出的指导,以及图像修复和基于文本的编辑技术。 该项目包括关于模型优化和开发的内容,涵盖概念微调和推理步骤的减少。它还提供了用于生成合成媒体的工作流,例如生成视频序列和将视觉频谱图转换为音频。 实践实现通过 PyTorch 代码示例和将模型权重及元数据发布到 Hugging Face Hub 的教程提供。
Provides a workflow for producing audio by generating visual spectrograms and converting them back into sound.
EchoMimic 是一个音频驱动的肖像动画框架和潜在扩散视频生成器。它通过将面部动作与音轨和运动驱动程序同步,将静态参考图像转换为动态的说话头像视频。 该系统作为一个混合运动合成引擎,结合了音频输入和姿势数据。它利用面部标志运动控制器来编辑定位标记,从而实现精确的同步和视频到视频的姿势迁移。 该管道通过潜在扩散和面部标志调节涵盖了图像到视频的动画。这允许由音频、姿势或两种引导源的组合驱动的肖像动画。
Implements neural encoders that translate speech features into facial expression parameters for synchronization.