awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

28 مستودعات

Awesome GitHub RepositoriesAudio Generation Models

Models specialized in processing and generating high-quality audio content.

Distinguishing note: Focuses on audio-specific model architectures.

Explore 28 awesome GitHub repositories matching artificial intelligence & ml · Audio Generation Models. Refine with filters or upvote what's useful.

Awesome Audio Generation Models GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • facebookresearch/audiocraftالصورة الرمزية لـ facebookresearch

    facebookresearch/audiocraft

    23,379عرض على GitHub↗

    Audiocraft is a deep learning audio library and machine learning framework designed for training, fine-tuning, and evaluating generative models for music and sound effects. It functions as a text-to-music generative model and a neural audio codec, providing the tools necessary to compress audio signals into discrete representations and synthesize high-fidelity waveforms from textual descriptions. The framework is distinguished by its ability to combine multiple conditioning signals, allowing for the generation of audio based on text prompts, melodic excerpts, or style-based audio clips. It al

    Generates subsequent audio content based on a provided starting audio clip to extend a sound sequence.

    Jupyter Notebook
    عرض على GitHub↗23,379
  • funaudiollm/cosyvoiceالصورة الرمزية لـ FunAudioLLM

    FunAudioLLM/CosyVoice

    21,673عرض على GitHub↗

    CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus

    Implements a neural synthesis architecture that modulates vocal attributes to produce speech with customizable emotional personas.

    Pythonaudio-generationcantonesechatbot
    عرض على GitHub↗21,673
  • nari-labs/diaالصورة الرمزية لـ nari-labs

    nari-labs/dia

    19,324عرض على GitHub↗

    Dia is a generative AI audio tool and text-to-speech synthesis engine designed for the production-ready deployment of machine learning models. It provides a framework for creating lifelike synthetic speech by conditioning generation on reference audio samples to replicate specific vocal characteristics, emotional tones, and delivery styles. The system distinguishes itself through its ability to perform custom voice cloning and precise control over audio output. Users can adjust generation parameters such as temperature and guidance scale to modify the pacing, creativity, and style of the synt

    Supports the injection of realistic nonverbal vocal expressions like laughter or gasps through specialized text markers.

    Pythonaiopen-weighttext-to-speech
    عرض على GitHub↗19,324
  • google-gemini/gemini-fullstack-langgraph-quickstartالصورة الرمزية لـ google-gemini

    google-gemini/gemini-fullstack-langgraph-quickstart

    18,217عرض على GitHub↗

    This project is an agentic workflow orchestrator designed for building and deploying autonomous systems that perform multi-step reasoning. It functions as a tool-augmented engine, enabling developers to chain model calls with external function execution to complete complex, user-defined tasks. By integrating large language models with persistent memory and stateful logic, the framework supports the creation of intelligent applications capable of independent operation. The platform distinguishes itself through graph-based state orchestration, which allows developers to define logic steps and t

    Provides high-quality, low-latency audio-to-audio models for real-time interaction.

    Jupyter Notebookgeminigemini-api
    عرض على GitHub↗18,217
  • google-gemini/cookbookالصورة الرمزية لـ google-gemini

    google-gemini/cookbook

    17,418عرض على GitHub↗

    The Gemini Cookbook is a comprehensive collection of implementation patterns, code samples, and development guides designed for building applications with Google Gemini models. It serves as a central resource for developers to integrate multimodal generative artificial intelligence into their software, providing the necessary frameworks to manage model interactions, stateful workflows, and structured data extraction. The repository distinguishes itself by offering specialized toolkits for autonomous agent orchestration, enabling the construction of agents that can execute code, browse the web

    Creates high-fidelity stereo music from text or image inputs, supporting custom lyrics and multi-language vocal performances.

    Jupyter Notebookgeminigemini-api
    عرض على GitHub↗17,418
  • sesameailabs/csmالصورة الرمزية لـ SesameAILabs

    SesameAILabs/csm

    14,669عرض على GitHub↗

    CSM is a conversational speech generation model and text-to-speech engine that converts text and audio inputs into synthetic speech. It utilizes a large language model architecture to predict and decode audio tokens for voice synthesis. The system functions as a zero-shot voice cloner, replicating specific speaker identities using short audio samples without requiring additional training. This enables precise control over speaker identity and the creation of synthetic speech that mimics a specific person. The model covers conversational speech synthesis and text-to-speech generation, transfo

    Utilizes a large language model architecture to predict and decode audio tokens for voice synthesis.

    Python
    عرض على GitHub↗14,669
  • alibaba/mnnالصورة الرمزية لـ alibaba

    alibaba/MNN

    14,242عرض على GitHub↗

    MNN is a high-performance inference engine and framework designed for on-device machine learning. It provides a comprehensive environment for executing, optimizing, and deploying neural network models directly on mobile and resource-constrained edge devices. The framework distinguishes itself through a robust model optimization toolkit that supports quantization, compression, and structural graph manipulation to minimize memory footprint and maximize execution speed. It features a modular architecture that abstracts hardware-specific backends, allowing models to run efficiently across diverse

    Generates audio waveform data from multimodal models for real-time streaming or file-based synthesis.

    C++armconvolutiondeep-learning
    عرض على GitHub↗14,242
  • winfredy/sadtalkerالصورة الرمزية لـ Winfredy

    Winfredy/SadTalker

    13,919عرض على GitHub↗

    SadTalker is a generative framework designed to synthesize expressive talking head videos from static portrait images. By mapping audio signals or text prompts to three-dimensional facial motion coefficients, the system synchronizes lip movements, facial expressions, and head orientation to create realistic digital character performances. The project distinguishes itself by decoupling identity from dynamic motion through latent space encoding, ensuring that the generated animations maintain visual fidelity to the source portrait. It supports comprehensive motion synthesis, including full-body

    Maps input audio signals to three-dimensional facial coefficients to synchronize lip movements and expressions with the source portrait.

    Python
    عرض على GitHub↗13,919
  • opentalker/sadtalkerالصورة الرمزية لـ OpenTalker

    OpenTalker/SadTalker

    13,895عرض على GitHub↗

    SadTalker is an audio-driven talking head generator that produces synchronized speaking videos from a single source image and an input audio file. The system utilizes a deep learning framework to map speech signals to facial motion data, enabling the creation of lifelike digital avatars and animated characters. The project distinguishes itself by employing a three-dimensional morphable model to translate audio features into precise facial landmarks and head pose parameters. It integrates latent diffusion motion synthesis to generate naturalistic head movements and uses expression-aware textur

    Extracts rhythmic and phonetic features from speech to drive the temporal evolution of facial expressions.

    Pythonaudio-driven-talking-facecvpr2023deep-fake
    عرض على GitHub↗13,895
  • facebookresearch/seamless_communicationالصورة الرمزية لـ facebookresearch

    facebookresearch/seamless_communication

    11,797عرض على GitHub↗

    This project is a multimodal translation framework and large language model capable of speech-to-speech, speech-to-text, and text-to-text translation across nearly 100 languages. It provides a real-time speech translation engine and a comprehensive toolkit for converting spoken audio between languages. The system is distinguished by its ability to preserve the original speaker's tone, pace, and prosody during translation. It utilizes a specialized on-device inference toolkit that converts model checkpoints into C-based libraries, enabling low-latency execution on mobile and edge hardware with

    Discovers pairs of audio segments across languages that share both the same meaning and expressivity.

    Jupyter Notebook
    عرض على GitHub↗11,797
  • aigc-audio/audiogptالصورة الرمزية لـ AIGC-Audio

    AIGC-Audio/AudioGPT

    10,174عرض على GitHub↗

    AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre

    Implements a system that generates soundscapes and audio clips from visual images or natural language descriptions.

    Pythonaudiogptmusic
    عرض على GitHub↗10,174
  • open-mmlab/amphionالصورة الرمزية لـ open-mmlab

    open-mmlab/Amphion

    9,844عرض على GitHub↗

    Amphion is an audio generation toolkit designed for the research and development of models that synthesize speech, music, and environmental sound effects. It provides a standardized framework for reproducible audio synthesis, incorporating a text-to-speech engine and a voice conversion framework. The project specializes in transforming audio identities, allowing for the modification of speaker accents and voice identities while preserving original rhythm and style. It also includes capabilities for singing voice synthesis and the generation of environmental soundscapes from text descriptions

    Produces high-quality audio waveforms from intermediate representations using specialized neural vocoders.

    Pythonaudio-generationaudio-synthesisaudioldm
    عرض على GitHub↗9,844
  • ml-explore/mlx-examplesالصورة الرمزية لـ ml-explore

    ml-explore/mlx-examples

    8,254عرض على GitHub↗

    This repository provides a collection of reference implementations and code examples for training and deploying machine learning models using the MLX framework. It serves as a practical guide for executing distributed training, fine-tuning large language models, converting model weights, and implementing multimodal generative workflows. The project distinguishes itself through specialized examples for local hardware execution, featuring weight quantization to reduce memory usage and low-rank adaptation for parameter-efficient fine-tuning. It also includes scripts for transforming external mod

    Implements generative workflows for producing text, images, audio, and video from mixed-modal inputs.

    Pythonmlx
    عرض على GitHub↗8,254
  • openai/jukeboxالصورة الرمزية لـ openai

    openai/jukebox

    8,039عرض على GitHub↗

    Jukebox is a generative audio model and AI music synthesis tool designed to create high-fidelity music samples and singing voices. It functions as a deep learning system that synthesizes raw audio conditioned on genre and artist metadata, utilizing a neural audio codec to convert raw audio into discrete codes for generative modeling and reconstruction. The system enables musical style steering and AI music composition by conditioning generation on specific artists, genres, and lyrics. It supports audio priming, allowing existing wave files to guide the creation of new musical sequences, and p

    Implements a generative audio model that predicts sequences of audio codes to create music.

    Pythonaudiogenerative-modelmusic
    عرض على GitHub↗8,039
  • humanaigc/emoالصورة الرمزية لـ HumanAIGC

    HumanAIGC/EMO

    7,616عرض على GitHub↗

    EMO هو نموذج ذكاء اصطناعي لتحريك الصور الشخصية ونموذج انتشار من الصوت إلى الفيديو، مصمم لتوليد مقاطع فيديو تعبيرية لأشخاص يتحدثون. يقوم بتحويل صورة شخصية ثابتة ومقطع صوتي إلى فيديو متزامن لشخص يتحدث. يركز النظام على تركيب الشخصيات الرقمية، وينتج حركات وجه وإشارات عاطفية عالية الدقة. يقوم بمزامنة حركات الشفاه وإيماءات الوجه لتطابق التسجيلات الصوتية لإنشاء رسوم متحركة واقعية للصور الشخصية. يستخدم إطار العمل عملية انتشار وآلية محاذاة متعددة الوسائط لضمان التوقيت بين الإشارات الصوتية ومعالم الوجه. كما يستخدم تكييف الصور القائم على المرجع للحفاظ على اتساق الهوية وطبقة اتساق زمنية لضمان سلاسة الحركة بين الإطارات.

    Generates high-fidelity human facial movements and emotional cues driven by audio signals.

    عرض على GitHub↗7,616
  • ibab/tensorflow-wavenetالصورة الرمزية لـ ibab

    ibab/tensorflow-wavenet

    5,432عرض على GitHub↗

    هذا المشروع عبارة عن تطبيق TensorFlow لشبكة عصبية لتوليد موجات صوتية خام. يعمل كنموذج لتوليف الكلام المشروط ينتج عينات صوتية اصطناعية باستخدام بنية شبكة عصبية تلافيفية موسعة (dilated convolutional neural network). يدعم النظام نمذجة الصوت المخصصة من خلال دمج التكييف العالمي والمعرفات الفئوية أثناء التدريب والتوليد. يسمح هذا للنموذج بتقليد متحدثين معينين أو خصائص صوتية مميزة لتطبيقات تحويل النص إلى كلام عصبية. يغطي إطار العمل توليف الصوت بالتعلم العميق، بما في ذلك معالجة مجموعات البيانات الصوتية، وتدريب النموذج من ملفات الموجات، وتوليد ملفات صوتية قابلة للتشغيل. يستخدم مكونات فنية مثل التلافيف السببية الموسعة، وضغط mu-law، ومخرجات softmax المكممة للتعامل مع التبعيات طويلة المدى في البيانات الصوتية.

    Provides a model capable of producing high-quality raw audio waveforms from trained weights.

    Python
    عرض على GitHub↗5,432
  • zejun-yang/aniportraitالصورة الرمزية لـ Zejun-Yang

    Zejun-Yang/AniPortrait

    5,020عرض على GitHub↗

    AniPortrait هو خط أنابيب لتوليف الفيديو بالذكاء الاصطناعي مصمم لإنشاء صور شخصية ناطقة واقعية ورسوم متحركة للوجه. يعمل كمولد للرؤوس المتحدثة ورسوم متحركة مدفوعة بالصوت تقوم بمزامنة حركات الشفاه، والتعبيرات، ووضعيات الرأس مع الكلام أو مصادر الفيديو المرجعية. يتضمن النظام أداة لنقل تعبيرات الوجه لإعادة تمثيل الحركات من فيديو مصدر على صورة مرجعية ثابتة. يستخدم نموذج انتشار كامن مع تكييف الصورة القائم على المرجع للحفاظ على الهوية البصرية والاتساق عبر الإطارات المولدة. يغطي خط الأنابيب تعيين الصوت إلى التعبير، والتحكم في الحركة الموجه بالوضعية، وتوليف الفيديو الواقعي. يدمج النظام زيادة دقة الإطارات (Upsampling) لتسريع عملية التوليد وتقليل وقت العرض الإجمالي.

    Implements neural encoders that extract features from audio to drive facial expression parameters.

    Python
    عرض على GitHub↗5,020
  • moonintheriver/diffsingerالصورة الرمزية لـ MoonInTheRiver

    MoonInTheRiver/DiffSinger

    4,804عرض على GitHub↗

    DiffSinger هو مركب صوتي للذكاء الاصطناعي ومولد صوت عصبي مصمم لإنتاج غناء وكلام عالي الدقة. يعمل كنظام تحويل النص إلى كلام وأداة تركيب صوت غنائي قائمة على الانتشار (Diffusion) تحول النص وطبقة الصوت إلى صوت مسموع. يستخدم النظام آلية انتشار ضحلة وتحسين ضوضاء تكراري لإنتاج عروض صوتية واقعية. ويدمج إضافات أخذ عينات متخصصة ومحلات عددية لتسريع الاستدلال وتقليل الوقت المطلوب لتوليد أصوات اصطناعية. يغطي المشروع النمذجة الصوتية، وتركيب مخطط ميل الطيفي (Mel-spectrogram)، وإعادة بناء الموكل العصبي (Neural vocoder) لتحويل النص إلى أشكال موجية صوتية في النطاق الزمني. كما يتضمن قدرات لتحسين الصوت الاصطناعي لتحسين الجودة الصوتية للتسجيلات.

    Produces high-fidelity audio waveforms and spectrograms for both singing and spoken language.

    Pythonaaai2022diffusion-modeldiffusion-speedup
    عرض على GitHub↗4,804
  • huggingface/diffusion-models-classالصورة الرمزية لـ huggingface

    huggingface/diffusion-models-class

    4,331عرض على GitHub↗

    This project is an educational course and collection of training materials focused on generative diffusion models. It provides a curriculum and practical guides for training, fine-tuning, and deploying models capable of synthesizing images, audio, and video. The material covers specific implementation strategies including noise-based synthesis, iterative refinement, and latent space compression. It provides instruction on guiding generative outputs through conditional synthesis and prompt adherence optimization, as well as techniques for image inpainting and text-based editing. The project i

    Provides a workflow for producing audio by generating visual spectrograms and converting them back into sound.

    Jupyter Notebook
    عرض على GitHub↗4,331
  • badtobest/echomimicالصورة الرمزية لـ BadToBest

    BadToBest/EchoMimic

    4,258عرض على GitHub↗

    EchoMimic is an audio-driven portrait animation framework and latent diffusion video generator. It transforms static reference images into dynamic talking head videos by synchronizing facial movements with audio tracks and motion drivers. The system functions as a hybrid motion synthesis engine that combines audio inputs and pose data. It utilizes a facial landmark motion controller to edit positioning markers, enabling precise synchronization and video-to-video pose transfer. The pipeline covers image-to-video animation through latent diffusion and facial landmark conditioning. This allows

    Implements neural encoders that translate speech features into facial expression parameters for synchronization.

    Python
    عرض على GitHub↗4,258
السابق12التالي
  1. Home
  2. Artificial Intelligence & ML
  3. Audio Generation Models

استكشف الوسوم الفرعية

  • Audio Prompt ContinuationCapabilities for extending an existing audio sequence based on a starting clip. **Distinct from Audio Generation Models:** Focuses on sequence continuation specifically rather than general generation
  • Audio Sample Reconstruction1 وسم فرعيProcesses that convert model samples into final audio waveforms with configurable formats. **Distinct from Audio Generation Models:** Focuses on the reconstruction of waveforms from model samples
  • Expressive Synthesis Models3 وسوم فرعيةNeural architectures that modulate vocal attributes and emotional personas during speech generation. **Distinct from Audio Generation Models:** Distinct from general audio generation: focuses specifically on expressive speech synthesis with emotional persona modulation.
  • Image-to-Audio Synthesis2 وسوم فرعيةSpecific neural models that translate visual content and context from images into corresponding audio. **Distinct from Audio Generation Models:** Narrowly focused on image-based triggers for audio generation, distinct from general audio generation models.
  • Integrity ValidationsVerification processes to ensure model outputs remain consistent with reference implementations. **Distinct from Audio Generation Models:** Focuses on correctness verification of processed waveforms rather than the generation process itself.
  • Melodic ConditioningGeneration of musical content guided by an existing audio melody. **Distinct from Audio Generation Models:** Specific to using melody as a guidance signal for generation
  • Multimodal GenerationAudio generation models that utilize non-audio inputs such as images or text to synthesize soundscapes. **Distinct from Audio Generation Models:** Focuses on cross-modal input mapping specifically, whereas Audio Generation Models is the broader category for any audio synthesis.