28 repositorios
Models specialized in processing and generating high-quality audio content.
Distinguishing note: Focuses on audio-specific model architectures.
Explore 28 awesome GitHub repositories matching artificial intelligence & ml · Audio Generation Models. Refine with filters or upvote what's useful.
Audiocraft is a deep learning audio library and machine learning framework designed for training, fine-tuning, and evaluating generative models for music and sound effects. It functions as a text-to-music generative model and a neural audio codec, providing the tools necessary to compress audio signals into discrete representations and synthesize high-fidelity waveforms from textual descriptions. The framework is distinguished by its ability to combine multiple conditioning signals, allowing for the generation of audio based on text prompts, melodic excerpts, or style-based audio clips. It al
Generates subsequent audio content based on a provided starting audio clip to extend a sound sequence.
CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus
Implements a neural synthesis architecture that modulates vocal attributes to produce speech with customizable emotional personas.
Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt
Supports the injection of realistic nonverbal vocal expressions like laughter or gasps through specialized text markers.
This project is an agentic workflow orchestrator designed for building and deploying autonomous systems that perform multi-step reasoning. It functions as a tool-augmented engine, enabling developers to chain model calls with external function execution to complete complex, user-defined tasks. By integrating large language models with persistent memory and stateful logic, the framework supports the creation of intelligent applications capable of independent operation. The platform distinguishes itself through graph-based state orchestration, which allows developers to define logic steps and t
Provides high-quality, low-latency audio-to-audio models for real-time interaction.
The Gemini Cookbook is a comprehensive collection of implementation patterns, code samples, and development guides designed for building applications with Google Gemini models. It serves as a central resource for developers to integrate multimodal generative artificial intelligence into their software, providing the necessary frameworks to manage model interactions, stateful workflows, and structured data extraction. The repository distinguishes itself by offering specialized toolkits for autonomous agent orchestration, enabling the construction of agents that can execute code, browse the web
Creates high-fidelity stereo music from text or image inputs, supporting custom lyrics and multi-language vocal performances.
CSM is a conversational speech generation model and text-to-speech engine that converts text and audio inputs into synthetic speech. It utilizes a large language model architecture to predict and decode audio tokens for voice synthesis. The system functions as a zero-shot voice cloner, replicating specific speaker identities using short audio samples without requiring additional training. This enables precise control over speaker identity and the creation of synthetic speech that mimics a specific person. The model covers conversational speech synthesis and text-to-speech generation, transfo
Utilizes a large language model architecture to predict and decode audio tokens for voice synthesis.
MNN is a high-performance inference engine and framework designed for on-device machine learning. It provides a comprehensive environment for executing, optimizing, and deploying neural network models directly on mobile and resource-constrained edge devices. The framework distinguishes itself through a robust model optimization toolkit that supports quantization, compression, and structural graph manipulation to minimize memory footprint and maximize execution speed. It features a modular architecture that abstracts hardware-specific backends, allowing models to run efficiently across diverse
Generates audio waveform data from multimodal models for real-time streaming or file-based synthesis.
SadTalker is a generative framework designed to synthesize expressive talking head videos from static portrait images. By mapping audio signals or text prompts to three-dimensional facial motion coefficients, the system synchronizes lip movements, facial expressions, and head orientation to create realistic digital character performances. The project distinguishes itself by decoupling identity from dynamic motion through latent space encoding, ensuring that the generated animations maintain visual fidelity to the source portrait. It supports comprehensive motion synthesis, including full-body
Maps input audio signals to three-dimensional facial coefficients to synchronize lip movements and expressions with the source portrait.
SadTalker is an audio-driven talking head generator that produces synchronized speaking videos from a single source image and an input audio file. The system utilizes a deep learning framework to map speech signals to facial motion data, enabling the creation of lifelike digital avatars and animated characters. The project distinguishes itself by employing a three-dimensional morphable model to translate audio features into precise facial landmarks and head pose parameters. It integrates latent diffusion motion synthesis to generate naturalistic head movements and uses expression-aware textur
Extracts rhythmic and phonetic features from speech to drive the temporal evolution of facial expressions.
This project is a multimodal translation framework and large language model capable of speech-to-speech, speech-to-text, and text-to-text translation across nearly 100 languages. It provides a real-time speech translation engine and a comprehensive toolkit for converting spoken audio between languages. The system is distinguished by its ability to preserve the original speaker's tone, pace, and prosody during translation. It utilizes a specialized on-device inference toolkit that converts model checkpoints into C-based libraries, enabling low-latency execution on mobile and edge hardware with
Discovers pairs of audio segments across languages that share both the same meaning and expressivity.
AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre
Implements a system that generates soundscapes and audio clips from visual images or natural language descriptions.
Amphion is an audio generation toolkit designed for the research and development of models that synthesize speech, music, and environmental sound effects. It provides a standardized framework for reproducible audio synthesis, incorporating a text-to-speech engine and a voice conversion framework. The project specializes in transforming audio identities, allowing for the modification of speaker accents and voice identities while preserving original rhythm and style. It also includes capabilities for singing voice synthesis and the generation of environmental soundscapes from text descriptions
Produces high-quality audio waveforms from intermediate representations using specialized neural vocoders.
This repository provides a collection of reference implementations and code examples for training and deploying machine learning models using the MLX framework. It serves as a practical guide for executing distributed training, fine-tuning large language models, converting model weights, and implementing multimodal generative workflows. The project distinguishes itself through specialized examples for local hardware execution, featuring weight quantization to reduce memory usage and low-rank adaptation for parameter-efficient fine-tuning. It also includes scripts for transforming external mod
Implements generative workflows for producing text, images, audio, and video from mixed-modal inputs.
Jukebox is a generative audio model and AI music synthesis tool designed to create high-fidelity music samples and singing voices. It functions as a deep learning system that synthesizes raw audio conditioned on genre and artist metadata, utilizing a neural audio codec to convert raw audio into discrete codes for generative modeling and reconstruction. The system enables musical style steering and AI music composition by conditioning generation on specific artists, genres, and lyrics. It supports audio priming, allowing existing wave files to guide the creation of new musical sequences, and p
Implements a generative audio model that predicts sequences of audio codes to create music.
EMO es un modelo de difusión de audio a video y animador de retratos por IA diseñado para generar videos expresivos de cabezas parlantes. Transforma una imagen de retrato estática y una pista de audio en un video sincronizado de una persona hablando. El sistema se centra en la síntesis de humanos digitales, produciendo movimientos faciales de alta fidelidad y señales emocionales. Sincroniza los movimientos de los labios y los gestos faciales con las grabaciones de voz para crear animaciones de retratos realistas. El framework utiliza un proceso de difusión y un mecanismo de alineación intermodal para asegurar la sincronización entre las señales de audio y los puntos de referencia visuales. Emplea un condicionamiento de imagen basado en referencias para mantener la consistencia de la identidad y una capa de consistencia temporal para asegurar un movimiento fluido entre fotogramas.
Generates high-fidelity human facial movements and emotional cues driven by audio signals.
Este proyecto es una implementación en TensorFlow de una red neuronal para la generación de formas de onda de audio crudas. Funciona como un modelo de síntesis de voz condicionado que produce muestras de audio sintéticas utilizando una arquitectura de red neuronal convolucional dilatada. El sistema admite el modelado de voz personalizado mediante la incorporación de condicionamiento global e identificadores categóricos durante el entrenamiento y la generación. Esto permite que el modelo imite hablantes específicos o características de audio distintas para aplicaciones de texto a voz neuronal. El framework cubre la síntesis de audio mediante aprendizaje profundo, incluyendo el procesamiento de datasets de audio, entrenamiento de modelos a partir de archivos de forma de onda y la generación de archivos de audio reproducibles. Utiliza componentes técnicos como convoluciones causales dilatadas, companding mu-law y salidas softmax cuantizadas para manejar dependencias de largo alcance en datos de audio.
Provides a model capable of producing high-quality raw audio waveforms from trained weights.
AniPortrait es un pipeline de síntesis de video por IA diseñado para generar retratos parlantes fotorrealistas y animaciones faciales. Funciona como un generador de cabezas parlantes y animador impulsado por audio que sincroniza los movimientos de los labios, las expresiones y las poses de la cabeza con fuentes de voz o video de referencia. El sistema incluye una herramienta de transferencia de expresiones faciales para recrear movimientos de un video fuente en una imagen de referencia estática. Utiliza un modelo de difusión latente con condicionamiento de imagen basado en referencia para mantener la identidad visual y la consistencia a través de los fotogramas generados. El pipeline cubre el mapeo de audio a expresión, el control de movimiento guiado por pose y la síntesis de video fotorrealista. Incorpora un upsampling de interpolación de fotogramas para acelerar el proceso de generación y reducir el tiempo total de renderizado.
Implements neural encoders that extract features from audio to drive facial expression parameters.
DiffSinger es un sintetizador vocal de IA y generador de audio neuronal diseñado para producir canto y habla de alta fidelidad. Funciona como un sistema de texto a voz y una herramienta de síntesis de voz cantada basada en difusión que transforma texto y tono en audio audible. El sistema utiliza un mecanismo de difusión superficial y refinamiento iterativo de ruido para generar interpretaciones vocales realistas. Incorpora plugins de muestreo especializados y solucionadores numéricos para acelerar la inferencia y reducir el tiempo requerido para generar voces sintéticas. El proyecto cubre el modelado acústico, la síntesis de mel-espectrogramas y la reconstrucción de vocoder neuronal para convertir texto en formas de onda de audio en el dominio del tiempo. También incluye capacidades para la mejora vocal sintética para mejorar la calidad sónica de las grabaciones.
Produces high-fidelity audio waveforms and spectrograms for both singing and spoken language.
This project is an educational course and collection of training materials focused on generative diffusion models. It provides a curriculum and practical guides for training, fine-tuning, and deploying models capable of synthesizing images, audio, and video. The material covers specific implementation strategies including noise-based synthesis, iterative refinement, and latent space compression. It provides instruction on guiding generative outputs through conditional synthesis and prompt adherence optimization, as well as techniques for image inpainting and text-based editing. The project i
Provides a workflow for producing audio by generating visual spectrograms and converting them back into sound.
EchoMimic is an audio-driven portrait animation framework and latent diffusion video generator. It transforms static reference images into dynamic talking head videos by synchronizing facial movements with audio tracks and motion drivers. The system functions as a hybrid motion synthesis engine that combines audio inputs and pose data. It utilizes a facial landmark motion controller to edit positioning markers, enabling precise synchronization and video-to-video pose transfer. The pipeline covers image-to-video animation through latent diffusion and facial landmark conditioning. This allows
Implements neural encoders that translate speech features into facial expression parameters for synchronization.