28 مستودعات
Models specialized in processing and generating high-quality audio content.
Distinguishing note: Focuses on audio-specific model architectures.
Explore 28 awesome GitHub repositories matching artificial intelligence & ml · Audio Generation Models. Refine with filters or upvote what's useful.
Audiocraft is a deep learning audio library and machine learning framework designed for training, fine-tuning, and evaluating generative models for music and sound effects. It functions as a text-to-music generative model and a neural audio codec, providing the tools necessary to compress audio signals into discrete representations and synthesize high-fidelity waveforms from textual descriptions. The framework is distinguished by its ability to combine multiple conditioning signals, allowing for the generation of audio based on text prompts, melodic excerpts, or style-based audio clips. It al
Generates subsequent audio content based on a provided starting audio clip to extend a sound sequence.
CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus
Implements a neural synthesis architecture that modulates vocal attributes to produce speech with customizable emotional personas.
Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt
Supports the injection of realistic nonverbal vocal expressions like laughter or gasps through specialized text markers.
This project is an agentic workflow orchestrator designed for building and deploying autonomous systems that perform multi-step reasoning. It functions as a tool-augmented engine, enabling developers to chain model calls with external function execution to complete complex, user-defined tasks. By integrating large language models with persistent memory and stateful logic, the framework supports the creation of intelligent applications capable of independent operation. The platform distinguishes itself through graph-based state orchestration, which allows developers to define logic steps and t
Provides high-quality, low-latency audio-to-audio models for real-time interaction.
The Gemini Cookbook is a comprehensive collection of implementation patterns, code samples, and development guides designed for building applications with Google Gemini models. It serves as a central resource for developers to integrate multimodal generative artificial intelligence into their software, providing the necessary frameworks to manage model interactions, stateful workflows, and structured data extraction. The repository distinguishes itself by offering specialized toolkits for autonomous agent orchestration, enabling the construction of agents that can execute code, browse the web
Creates high-fidelity stereo music from text or image inputs, supporting custom lyrics and multi-language vocal performances.
CSM is a conversational speech generation model and text-to-speech engine that converts text and audio inputs into synthetic speech. It utilizes a large language model architecture to predict and decode audio tokens for voice synthesis. The system functions as a zero-shot voice cloner, replicating specific speaker identities using short audio samples without requiring additional training. This enables precise control over speaker identity and the creation of synthetic speech that mimics a specific person. The model covers conversational speech synthesis and text-to-speech generation, transfo
Utilizes a large language model architecture to predict and decode audio tokens for voice synthesis.
MNN is a high-performance inference engine and framework designed for on-device machine learning. It provides a comprehensive environment for executing, optimizing, and deploying neural network models directly on mobile and resource-constrained edge devices. The framework distinguishes itself through a robust model optimization toolkit that supports quantization, compression, and structural graph manipulation to minimize memory footprint and maximize execution speed. It features a modular architecture that abstracts hardware-specific backends, allowing models to run efficiently across diverse
Generates audio waveform data from multimodal models for real-time streaming or file-based synthesis.
SadTalker is a generative framework designed to synthesize expressive talking head videos from static portrait images. By mapping audio signals or text prompts to three-dimensional facial motion coefficients, the system synchronizes lip movements, facial expressions, and head orientation to create realistic digital character performances. The project distinguishes itself by decoupling identity from dynamic motion through latent space encoding, ensuring that the generated animations maintain visual fidelity to the source portrait. It supports comprehensive motion synthesis, including full-body
Maps input audio signals to three-dimensional facial coefficients to synchronize lip movements and expressions with the source portrait.
SadTalker is an audio-driven talking head generator that produces synchronized speaking videos from a single source image and an input audio file. The system utilizes a deep learning framework to map speech signals to facial motion data, enabling the creation of lifelike digital avatars and animated characters. The project distinguishes itself by employing a three-dimensional morphable model to translate audio features into precise facial landmarks and head pose parameters. It integrates latent diffusion motion synthesis to generate naturalistic head movements and uses expression-aware textur
Extracts rhythmic and phonetic features from speech to drive the temporal evolution of facial expressions.
This project is a multimodal translation framework and large language model capable of speech-to-speech, speech-to-text, and text-to-text translation across nearly 100 languages. It provides a real-time speech translation engine and a comprehensive toolkit for converting spoken audio between languages. The system is distinguished by its ability to preserve the original speaker's tone, pace, and prosody during translation. It utilizes a specialized on-device inference toolkit that converts model checkpoints into C-based libraries, enabling low-latency execution on mobile and edge hardware with
Discovers pairs of audio segments across languages that share both the same meaning and expressivity.
AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre
Implements a system that generates soundscapes and audio clips from visual images or natural language descriptions.
Amphion is an audio generation toolkit designed for the research and development of models that synthesize speech, music, and environmental sound effects. It provides a standardized framework for reproducible audio synthesis, incorporating a text-to-speech engine and a voice conversion framework. The project specializes in transforming audio identities, allowing for the modification of speaker accents and voice identities while preserving original rhythm and style. It also includes capabilities for singing voice synthesis and the generation of environmental soundscapes from text descriptions
Produces high-quality audio waveforms from intermediate representations using specialized neural vocoders.
This repository provides a collection of reference implementations and code examples for training and deploying machine learning models using the MLX framework. It serves as a practical guide for executing distributed training, fine-tuning large language models, converting model weights, and implementing multimodal generative workflows. The project distinguishes itself through specialized examples for local hardware execution, featuring weight quantization to reduce memory usage and low-rank adaptation for parameter-efficient fine-tuning. It also includes scripts for transforming external mod
Implements generative workflows for producing text, images, audio, and video from mixed-modal inputs.
Jukebox is a generative audio model and AI music synthesis tool designed to create high-fidelity music samples and singing voices. It functions as a deep learning system that synthesizes raw audio conditioned on genre and artist metadata, utilizing a neural audio codec to convert raw audio into discrete codes for generative modeling and reconstruction. The system enables musical style steering and AI music composition by conditioning generation on specific artists, genres, and lyrics. It supports audio priming, allowing existing wave files to guide the creation of new musical sequences, and p
Implements a generative audio model that predicts sequences of audio codes to create music.
EMO هو نموذج ذكاء اصطناعي لتحريك الصور الشخصية ونموذج انتشار من الصوت إلى الفيديو، مصمم لتوليد مقاطع فيديو تعبيرية لأشخاص يتحدثون. يقوم بتحويل صورة شخصية ثابتة ومقطع صوتي إلى فيديو متزامن لشخص يتحدث. يركز النظام على تركيب الشخصيات الرقمية، وينتج حركات وجه وإشارات عاطفية عالية الدقة. يقوم بمزامنة حركات الشفاه وإيماءات الوجه لتطابق التسجيلات الصوتية لإنشاء رسوم متحركة واقعية للصور الشخصية. يستخدم إطار العمل عملية انتشار وآلية محاذاة متعددة الوسائط لضمان التوقيت بين الإشارات الصوتية ومعالم الوجه. كما يستخدم تكييف الصور القائم على المرجع للحفاظ على اتساق الهوية وطبقة اتساق زمنية لضمان سلاسة الحركة بين الإطارات.
Generates high-fidelity human facial movements and emotional cues driven by audio signals.
هذا المشروع عبارة عن تطبيق TensorFlow لشبكة عصبية لتوليد موجات صوتية خام. يعمل كنموذج لتوليف الكلام المشروط ينتج عينات صوتية اصطناعية باستخدام بنية شبكة عصبية تلافيفية موسعة (dilated convolutional neural network). يدعم النظام نمذجة الصوت المخصصة من خلال دمج التكييف العالمي والمعرفات الفئوية أثناء التدريب والتوليد. يسمح هذا للنموذج بتقليد متحدثين معينين أو خصائص صوتية مميزة لتطبيقات تحويل النص إلى كلام عصبية. يغطي إطار العمل توليف الصوت بالتعلم العميق، بما في ذلك معالجة مجموعات البيانات الصوتية، وتدريب النموذج من ملفات الموجات، وتوليد ملفات صوتية قابلة للتشغيل. يستخدم مكونات فنية مثل التلافيف السببية الموسعة، وضغط mu-law، ومخرجات softmax المكممة للتعامل مع التبعيات طويلة المدى في البيانات الصوتية.
Provides a model capable of producing high-quality raw audio waveforms from trained weights.
AniPortrait هو خط أنابيب لتوليف الفيديو بالذكاء الاصطناعي مصمم لإنشاء صور شخصية ناطقة واقعية ورسوم متحركة للوجه. يعمل كمولد للرؤوس المتحدثة ورسوم متحركة مدفوعة بالصوت تقوم بمزامنة حركات الشفاه، والتعبيرات، ووضعيات الرأس مع الكلام أو مصادر الفيديو المرجعية. يتضمن النظام أداة لنقل تعبيرات الوجه لإعادة تمثيل الحركات من فيديو مصدر على صورة مرجعية ثابتة. يستخدم نموذج انتشار كامن مع تكييف الصورة القائم على المرجع للحفاظ على الهوية البصرية والاتساق عبر الإطارات المولدة. يغطي خط الأنابيب تعيين الصوت إلى التعبير، والتحكم في الحركة الموجه بالوضعية، وتوليف الفيديو الواقعي. يدمج النظام زيادة دقة الإطارات (Upsampling) لتسريع عملية التوليد وتقليل وقت العرض الإجمالي.
Implements neural encoders that extract features from audio to drive facial expression parameters.
DiffSinger هو مركب صوتي للذكاء الاصطناعي ومولد صوت عصبي مصمم لإنتاج غناء وكلام عالي الدقة. يعمل كنظام تحويل النص إلى كلام وأداة تركيب صوت غنائي قائمة على الانتشار (Diffusion) تحول النص وطبقة الصوت إلى صوت مسموع. يستخدم النظام آلية انتشار ضحلة وتحسين ضوضاء تكراري لإنتاج عروض صوتية واقعية. ويدمج إضافات أخذ عينات متخصصة ومحلات عددية لتسريع الاستدلال وتقليل الوقت المطلوب لتوليد أصوات اصطناعية. يغطي المشروع النمذجة الصوتية، وتركيب مخطط ميل الطيفي (Mel-spectrogram)، وإعادة بناء الموكل العصبي (Neural vocoder) لتحويل النص إلى أشكال موجية صوتية في النطاق الزمني. كما يتضمن قدرات لتحسين الصوت الاصطناعي لتحسين الجودة الصوتية للتسجيلات.
Produces high-fidelity audio waveforms and spectrograms for both singing and spoken language.
This project is an educational course and collection of training materials focused on generative diffusion models. It provides a curriculum and practical guides for training, fine-tuning, and deploying models capable of synthesizing images, audio, and video. The material covers specific implementation strategies including noise-based synthesis, iterative refinement, and latent space compression. It provides instruction on guiding generative outputs through conditional synthesis and prompt adherence optimization, as well as techniques for image inpainting and text-based editing. The project i
Provides a workflow for producing audio by generating visual spectrograms and converting them back into sound.
EchoMimic is an audio-driven portrait animation framework and latent diffusion video generator. It transforms static reference images into dynamic talking head videos by synchronizing facial movements with audio tracks and motion drivers. The system functions as a hybrid motion synthesis engine that combines audio inputs and pose data. It utilizes a facial landmark motion controller to edit positioning markers, enabling precise synchronization and video-to-video pose transfer. The pipeline covers image-to-video animation through latent diffusion and facial landmark conditioning. This allows
Implements neural encoders that translate speech features into facial expression parameters for synchronization.