28 Repos
Models specialized in processing and generating high-quality audio content.
Distinguishing note: Focuses on audio-specific model architectures.
Explore 28 awesome GitHub repositories matching artificial intelligence & ml · Audio Generation Models. Refine with filters or upvote what's useful.
Audiocraft is a deep learning audio library and machine learning framework designed for training, fine-tuning, and evaluating generative models for music and sound effects. It functions as a text-to-music generative model and a neural audio codec, providing the tools necessary to compress audio signals into discrete representations and synthesize high-fidelity waveforms from textual descriptions. The framework is distinguished by its ability to combine multiple conditioning signals, allowing for the generation of audio based on text prompts, melodic excerpts, or style-based audio clips. It al
Generates subsequent audio content based on a provided starting audio clip to extend a sound sequence.
CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus
Implements a neural synthesis architecture that modulates vocal attributes to produce speech with customizable emotional personas.
Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt
Supports the injection of realistic nonverbal vocal expressions like laughter or gasps through specialized text markers.
This project is an agentic workflow orchestrator designed for building and deploying autonomous systems that perform multi-step reasoning. It functions as a tool-augmented engine, enabling developers to chain model calls with external function execution to complete complex, user-defined tasks. By integrating large language models with persistent memory and stateful logic, the framework supports the creation of intelligent applications capable of independent operation. The platform distinguishes itself through graph-based state orchestration, which allows developers to define logic steps and t
Provides high-quality, low-latency audio-to-audio models for real-time interaction.
The Gemini Cookbook is a comprehensive collection of implementation patterns, code samples, and development guides designed for building applications with Google Gemini models. It serves as a central resource for developers to integrate multimodal generative artificial intelligence into their software, providing the necessary frameworks to manage model interactions, stateful workflows, and structured data extraction. The repository distinguishes itself by offering specialized toolkits for autonomous agent orchestration, enabling the construction of agents that can execute code, browse the web
Creates high-fidelity stereo music from text or image inputs, supporting custom lyrics and multi-language vocal performances.
CSM is a conversational speech generation model and text-to-speech engine that converts text and audio inputs into synthetic speech. It utilizes a large language model architecture to predict and decode audio tokens for voice synthesis. The system functions as a zero-shot voice cloner, replicating specific speaker identities using short audio samples without requiring additional training. This enables precise control over speaker identity and the creation of synthetic speech that mimics a specific person. The model covers conversational speech synthesis and text-to-speech generation, transfo
Utilizes a large language model architecture to predict and decode audio tokens for voice synthesis.
MNN is a high-performance inference engine and framework designed for on-device machine learning. It provides a comprehensive environment for executing, optimizing, and deploying neural network models directly on mobile and resource-constrained edge devices. The framework distinguishes itself through a robust model optimization toolkit that supports quantization, compression, and structural graph manipulation to minimize memory footprint and maximize execution speed. It features a modular architecture that abstracts hardware-specific backends, allowing models to run efficiently across diverse
Generates audio waveform data from multimodal models for real-time streaming or file-based synthesis.
SadTalker is a generative framework designed to synthesize expressive talking head videos from static portrait images. By mapping audio signals or text prompts to three-dimensional facial motion coefficients, the system synchronizes lip movements, facial expressions, and head orientation to create realistic digital character performances. The project distinguishes itself by decoupling identity from dynamic motion through latent space encoding, ensuring that the generated animations maintain visual fidelity to the source portrait. It supports comprehensive motion synthesis, including full-body
Maps input audio signals to three-dimensional facial coefficients to synchronize lip movements and expressions with the source portrait.
SadTalker is an audio-driven talking head generator that produces synchronized speaking videos from a single source image and an input audio file. The system utilizes a deep learning framework to map speech signals to facial motion data, enabling the creation of lifelike digital avatars and animated characters. The project distinguishes itself by employing a three-dimensional morphable model to translate audio features into precise facial landmarks and head pose parameters. It integrates latent diffusion motion synthesis to generate naturalistic head movements and uses expression-aware textur
Extracts rhythmic and phonetic features from speech to drive the temporal evolution of facial expressions.
This project is a multimodal translation framework and large language model capable of speech-to-speech, speech-to-text, and text-to-text translation across nearly 100 languages. It provides a real-time speech translation engine and a comprehensive toolkit for converting spoken audio between languages. The system is distinguished by its ability to preserve the original speaker's tone, pace, and prosody during translation. It utilizes a specialized on-device inference toolkit that converts model checkpoints into C-based libraries, enabling low-latency execution on mobile and edge hardware with
Discovers pairs of audio segments across languages that share both the same meaning and expressivity.
AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre
Implements a system that generates soundscapes and audio clips from visual images or natural language descriptions.
Amphion is an audio generation toolkit designed for the research and development of models that synthesize speech, music, and environmental sound effects. It provides a standardized framework for reproducible audio synthesis, incorporating a text-to-speech engine and a voice conversion framework. The project specializes in transforming audio identities, allowing for the modification of speaker accents and voice identities while preserving original rhythm and style. It also includes capabilities for singing voice synthesis and the generation of environmental soundscapes from text descriptions
Produces high-quality audio waveforms from intermediate representations using specialized neural vocoders.
This repository provides a collection of reference implementations and code examples for training and deploying machine learning models using the MLX framework. It serves as a practical guide for executing distributed training, fine-tuning large language models, converting model weights, and implementing multimodal generative workflows. The project distinguishes itself through specialized examples for local hardware execution, featuring weight quantization to reduce memory usage and low-rank adaptation for parameter-efficient fine-tuning. It also includes scripts for transforming external mod
Implements generative workflows for producing text, images, audio, and video from mixed-modal inputs.
Jukebox is a generative audio model and AI music synthesis tool designed to create high-fidelity music samples and singing voices. It functions as a deep learning system that synthesizes raw audio conditioned on genre and artist metadata, utilizing a neural audio codec to convert raw audio into discrete codes for generative modeling and reconstruction. The system enables musical style steering and AI music composition by conditioning generation on specific artists, genres, and lyrics. It supports audio priming, allowing existing wave files to guide the creation of new musical sequences, and p
Implements a generative audio model that predicts sequences of audio codes to create music.
EMO ist ein KI-Porträt-Animator und Audio-zu-Video-Diffusionsmodell, das entwickelt wurde, um ausdrucksstarke Talking-Head-Videos zu generieren. Es verwandelt ein einzelnes statisches Porträtbild und eine Audiospur in ein synchronisiertes Video einer sprechenden Person. Das System konzentriert sich auf die Synthese digitaler Menschen und erzeugt hochauflösende Gesichtsbewegungen und emotionale Signale. Es synchronisiert Lippenbewegungen und Gesichtsausdrücke mit gesprochenen Sprachaufnahmen, um realistische Porträt-Animationen zu erstellen. Das Framework nutzt einen Diffusionsprozess und einen Cross-Modal-Alignment-Mechanismus, um das Timing zwischen Audiosignalen und visuellen Landmarks sicherzustellen. Es verwendet referenzbasierte Bildkonditionierung, um die Identitätskonsistenz zu wahren, sowie eine zeitliche Konsistenzschicht, um flüssige Bewegungen zwischen den Frames zu gewährleisten.
Generates high-fidelity human facial movements and emotional cues driven by audio signals.
Dieses Projekt ist eine TensorFlow-Implementierung eines neuronalen Netzwerks für die Generierung von rohen Audiowellenformen. Es fungiert als konditioniertes Sprachsynthese-Modell, das synthetische Audio-Samples unter Verwendung einer Architektur mit dilatierten konvolutionalen neuronalen Netzwerken erzeugt. Das System unterstützt benutzerdefiniertes Voice-Modeling durch die Einbeziehung globaler Konditionierung und kategorialer Identifikatoren während des Trainings und der Generierung. Dies ermöglicht es dem Modell, spezifische Sprecher oder ausgeprägte Audiocharakteristika für neuronale Text-to-Speech-Anwendungen nachzuahmen. Das Framework deckt Deep-Learning-Audiosynthese ab, einschließlich der Verarbeitung von Audiodatensätzen, des Modelltrainings aus Wellenformdateien und der Generierung abspielbarer Audiodateien. Es nutzt technische Komponenten wie dilatierte kausale Konvolutionen, Mu-Law-Kompandierung und quantisierte Softmax-Ausgaben, um Langzeitabhängigkeiten in Audiodaten zu verarbeiten.
Provides a model capable of producing high-quality raw audio waveforms from trained weights.
AniPortrait is an AI video synthesis pipeline designed to generate photorealistic speaking portraits and facial animations. It functions as a talking head generator and audio-driven animator that synchronizes lip movements, expressions, and head poses to speech or reference video sources. The system includes a facial expression transfer tool for reenacting movements from a source video onto a static reference image. It utilizes a latent diffusion model with reference-based image conditioning to maintain visual identity and consistency across generated frames. The pipeline covers audio-to-exp
Implements neural encoders that extract features from audio to drive facial expression parameters.
DiffSinger ist ein KI-Gesangssynthesizer und neuronaler Audiogenerator, der darauf ausgelegt ist, hochqualitativen Gesang und Sprache zu produzieren. Er fungiert als Text-to-Speech-System und als diffusionsbasiertes Tool zur Synthese von Gesangsstimmen, das Text und Tonhöhe in hörbares Audio transformiert. Das System nutzt einen flachen Diffusionsmechanismus und iterative Rauschverfeinerung, um realistische Gesangsdarbietungen zu generieren. Es integriert spezialisierte Sampling-Plugins und numerische Löser, um die Inferenz zu beschleunigen und die Zeit zu reduzieren, die zur Generierung synthetischer Stimmen erforderlich ist. Das Projekt deckt akustische Modellierung, Mel-Spektrogramm-Synthese und neuronale Vocoder-Rekonstruktion ab, um Text in Zeitbereichs-Audio-Wellenformen zu konvertieren. Es enthält zudem Funktionen zur synthetischen Stimmverbesserung, um die klangliche Qualität von Aufnahmen zu steigern.
Produces high-fidelity audio waveforms and spectrograms for both singing and spoken language.
This project is an educational course and collection of training materials focused on generative diffusion models. It provides a curriculum and practical guides for training, fine-tuning, and deploying models capable of synthesizing images, audio, and video. The material covers specific implementation strategies including noise-based synthesis, iterative refinement, and latent space compression. It provides instruction on guiding generative outputs through conditional synthesis and prompt adherence optimization, as well as techniques for image inpainting and text-based editing. The project i
Provides a workflow for producing audio by generating visual spectrograms and converting them back into sound.
EchoMimic is an audio-driven portrait animation framework and latent diffusion video generator. It transforms static reference images into dynamic talking head videos by synchronizing facial movements with audio tracks and motion drivers. The system functions as a hybrid motion synthesis engine that combines audio inputs and pose data. It utilizes a facial landmark motion controller to edit positioning markers, enabling precise synchronization and video-to-video pose transfer. The pipeline covers image-to-video animation through latent diffusion and facial landmark conditioning. This allows
Implements neural encoders that translate speech features into facial expression parameters for synchronization.