awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

47 个仓库

Awesome GitHub RepositoriesAudio Synthesis

Systems that generate artificial audio signals, including advanced neural vocoders for voice and sound synthesis.

Explore 47 awesome GitHub repositories matching graphics & multimedia · Audio Synthesis. Refine with filters or upvote what's useful.

Awesome Audio Synthesis GitHub Repositories

用 AI 发现最棒的仓库。我们将通过 AI 为您搜索最匹配的仓库。
  • corentinj/real-time-voice-cloningCorentinJ 的头像

    CorentinJ/Real-Time-Voice-Cloning

    59,918在 GitHub 上查看↗

    This project is a neural text-to-speech engine and voice cloning toolkit designed to generate synthetic speech that mimics the vocal characteristics of a target speaker. It functions as a real-time audio synthesizer, utilizing a deep learning pipeline to convert written text into high-fidelity speech output with minimal latency. The system employs a transfer learning framework that leverages pre-trained speaker verification models to adapt synthesis to new, unseen vocal identities. By using an encoder-based speaker embedding process, the toolkit maps variable-length audio samples into a laten

    Synthesizes high-fidelity audio waveforms from spectral representations using models optimized for rapid inference.

    Pythondeep-learningpythonpytorch
    在 GitHub 上查看↗59,918
  • rvc-boss/gpt-sovitsRVC-Boss 的头像

    RVC-Boss/GPT-SoVITS

    58,724在 GitHub 上查看↗

    GPT-SoVITS is a text-to-speech synthesis engine and voice cloning toolkit designed for generating natural-sounding human speech. It functions as a neural audio processing pipeline that maps input text to high-fidelity audio waveforms, utilizing conditional variational autoencoders and flow-based decoders to ensure expressive output. The platform distinguishes itself through its ability to perform few-shot voice cloning and cross-lingual speech generation, allowing users to maintain a specific speaker's vocal identity and emotional delivery across multiple languages. By employing cross-modal l

    Transforms generated spectral data into high-fidelity time-domain audio waveforms using specialized neural models.

    Pythontext-to-speechttsvits
    在 GitHub 上查看↗58,724
  • microsoft/vibevoicemicrosoft 的头像

    microsoft/VibeVoice

    49,394在 GitHub 上查看↗

    VibeVoice is a generative artificial intelligence platform designed for text-to-speech synthesis. It functions as a neural audio generation framework that converts written text into natural-sounding spoken audio, specifically engineered to maintain consistent vocal characteristics and narrative prosody across extended passages of content. The system distinguishes itself through its ability to generate long-form conversational speech while preserving speaker identity and linguistic content. By utilizing latent space disentanglement, the model separates speaker traits from the input text, allow

    Converts high-level acoustic features into raw waveform audio using deep generative neural vocoders.

    Python
    在 GitHub 上查看↗49,394
  • coqui-ai/ttscoqui-ai 的头像

    coqui-ai/TTS

    45,568在 GitHub 上查看↗

    这是一个深度学习文本转语音工具包,用于训练和部署神经语音合成模型。它提供了一个完整的框架,用于将书面文本转换为口语音频,利用神经声码器将合成的频谱图转换为高保真音频波形。 该工具包包括一个语音克隆系统,通过从短音频样本中提取说话人嵌入来复制特定的人声。它还支持多说话人音频合成,允许使用专门的模型架构生成不同声线身份的语音。 该系统涵盖了完整的语音合成流水线,包括语音数据集整理工具、带有性能跟踪的模型自定义训练,以及用于音频生成的命令行界面。对于网络访问,它提供了一个自托管的 HTTP 服务器,将语音合成模型部署为 API。

    Includes neural vocoders that transform synthesized spectrograms into high-fidelity time-domain audio waveforms.

    Pythondeep-learningglow-ttshifigan
    在 GitHub 上查看↗45,568
  • babysor/mockingbirdbabysor 的头像

    babysor/MockingBird

    36,903在 GitHub 上查看↗

    MockingBird is an AI voice cloning tool and text-to-speech system designed to generate synthetic speech. It functions as a voice synthesis trainer for building custom models from audio datasets, a command-line generator for producing audio files, and a text-to-speech server for remote application integration. The project specializes in real-time voice cloning, which extracts vocal characteristics from short audio samples to mimic a target speaker's unique timbre. It utilizes reference-driven audio synthesis to condition pre-trained models on specific audio samples, allowing for the generation

    Generates arbitrary speech conditioned on a specific audio sample to maintain voice identity.

    Pythonaideep-learningpytorch
    在 GitHub 上查看↗36,903
  • rvc-project/retrieval-based-voice-conversion-webuiRVC-Project 的头像

    RVC-Project/Retrieval-based-Voice-Conversion-WebUI

    36,025在 GitHub 上查看↗

    This project is a comprehensive software suite for voice synthesis and model management, providing a framework for training custom acoustic models and performing voice conversion. It utilizes deep-learning-based acoustic modeling to map source audio characteristics to target voice identities, enabling the transformation of input audio into specific vocal profiles. The system distinguishes itself through a feature-retrieval-based inference mechanism, which employs vector index files to perform nearest-neighbor searches on acoustic features for high-fidelity timbre matching. Users can manage th

    Synthesizes vocal characteristics by applying learned voice models to input audio sources.

    Pythonaudio-analysischangeconversational-ai
    在 GitHub 上查看↗36,025
  • svc-develop-team/so-vits-svcsvc-develop-team 的头像

    svc-develop-team/so-vits-svc

    28,097在 GitHub 上查看↗

    This project is a singing voice conversion tool based on VITS generative modeling. It transforms the identity of a singing voice to a target speaker while preserving the original melody, lyrics, and intonation. The system distinguishes itself through hybrid voice synthesis, allowing for the blending of multiple speaker identities via linear model interpolation. It utilizes cluster-based feature retrieval to increase target voice similarity and employs a diffusion probabilistic model as a post-processor to remove electronic artifacts and improve vocal clarity. The software covers a broad rang

    Blends multiple speaker models to create hybrid voice identities through linear interpolation.

    Python
    在 GitHub 上查看↗28,097
  • resemble-ai/chatterboxresemble-ai 的头像

    resemble-ai/chatterbox

    22,751在 GitHub 上查看↗

    Chatterbox is a comprehensive machine learning platform designed for multilingual speech synthesis and real-time audio generation. It functions as an engine that converts text into natural-sounding speech, capable of replicating specific human vocal characteristics and emotional expressions from short audio samples. The platform distinguishes itself through advanced control over the synthesis process, allowing for the manipulation of emotional intensity and the injection of non-verbal vocalizations such as laughter or coughing. It is engineered for low-latency performance, utilizing an optimi

    Transforms generated spectral data into high-fidelity time-domain audio waveforms.

    Python
    在 GitHub 上查看↗22,751
  • funaudiollm/cosyvoiceFunAudioLLM 的头像

    FunAudioLLM/CosyVoice

    21,673在 GitHub 上查看↗

    CosyVoice is a speech synthesis framework that utilizes large language models to generate expressive, multilingual audio. The system functions as an audio generation engine capable of producing natural-sounding speech across multiple languages while preserving regional dialects and specific emotional tones. The platform distinguishes itself through its zero-shot voice cloning capabilities, which allow for the creation of synthetic voice profiles from short audio samples without requiring additional model training. It provides fine-grained control over vocal attributes, enabling users to adjus

    Converts raw acoustic tokens into high-fidelity waveforms using deep learning models.

    Pythonaudio-generationcantonesechatbot
    在 GitHub 上查看↗21,673
  • magenta/magentamagenta 的头像

    magenta/magenta

    19,778在 GitHub 上查看↗

    Magenta is a comprehensive toolkit for training, synthesizing, and performing music through neural models and hardware-integrated engines. It functions as a machine learning framework that enables the generation, manipulation, and real-time performance of audio, providing the structural foundations for musical intelligence through hierarchical sequence modeling and symbolic processing. The project distinguishes itself by enabling real-time, low-latency neural audio synthesis that can be integrated directly into professional digital audio workstations. It supports interactive musical jamming a

    Transforms learned feature representations into raw waveforms using deep learning models.

    Python
    在 GitHub 上查看↗19,778
  • index-tts/index-ttsindex-tts 的头像

    index-tts/index-tts

    18,851在 GitHub 上查看↗

    Index-tts is a neural audio generation engine designed to convert written text into high-fidelity human speech. By utilizing deep learning models and phoneme-based sequence modeling, the system transforms text into natural-sounding audio waveforms suitable for a variety of accessibility and media applications. The platform functions as a server-side inference pipeline that provides a programmatic interface for integrating voice generation into external applications. It distinguishes itself through asynchronous audio streaming, which buffers and delivers generated speech chunks in real time to

    Transforms linguistic feature representations into high-fidelity raw audio waveforms using deep learning models.

    Pythonbigvgancross-lingualindextts
    在 GitHub 上查看↗18,851
  • kitao/pyxelkitao 的头像

    kitao/pyxel

    17,547在 GitHub 上查看↗

    Pyxel is a Python-based retro game engine designed for creating pixel-art applications with vintage aesthetics. It functions as a pixel art framework that emulates classic hardware limitations through fixed-resolution displays, constrained color palettes, and a 2D tilemap engine for rendering scaled sprites and grid-based game worlds. The engine features a dedicated MML audio synthesizer that uses Macro Language notation to compose and play back retro-style sound sequences. It also includes capabilities for seed-based background music generation and the playback of raw PCM audio data for cust

    Includes an audio synthesis system that generates wave forms in real-time from MML notation.

    Rust
    在 GitHub 上查看↗17,547
  • huanshere/videolingoHuanshere 的头像

    Huanshere/VideoLingo

    17,498在 GitHub 上查看↗

    VideoLingo is an automated video localization suite designed to transcribe, translate, and dub video content. It functions as a translation pipeline that utilizes large language models to convert spoken audio into precise text segments and translate them into multiple languages. The system differentiates itself through a multi-step translation refinement process and a specialized natural language processing utility that segments text into single-line captions meeting broadcast standards. It also integrates synthetic voiceover generation to replace or augment original audio tracks. The projec

    Generates artificial voiceovers based on translated subtitles to replace original audio.

    Pythonai-translationdubbinglocalization
    在 GitHub 上查看↗17,498
  • musescore/musescoremusescore 的头像

    musescore/MuseScore

    14,732在 GitHub 上查看↗

    MuseScore is a professional music notation application designed for composing, arranging, and engraving musical scores. It provides a graphical interface that renders notation in real-time, allowing users to create and edit complex musical arrangements with immediate visual feedback. The software distinguishes itself through a robust document-object model that manages the relationships between notes, staves, and layout formatting. It supports the standard markup language for music interchange, ensuring that scores can be shared across different notation platforms. Additionally, the applicatio

    Processes musical data through a low-latency pipeline to generate high-fidelity sound output based on loaded soundfont samples.

    C++cppmusescoremusic-notation
    在 GitHub 上查看↗14,732
  • hammerspoon/hammerspoonHammerspoon 的头像

    Hammerspoon/hammerspoon

    14,497在 GitHub 上查看↗

    Hammerspoon is a programmable automation engine for macOS that enables deep system-level control through a Lua scripting environment. By bridging high-level scripts with native Objective-C APIs, it allows users to interact with the operating system's accessibility tree, intercept hardware input streams, and manage the lifecycle of running applications. The project distinguishes itself through an event-driven architecture that registers asynchronous hooks for system notifications and hardware events. This allows for real-time automation, such as remapping keyboard and mouse inputs, managing wi

    Controls internal MIDI audio synthesis directly from automation scripts.

    Objective-Cautomationhammerspoonirc
    在 GitHub 上查看↗14,497
  • paddlepaddle/paddlespeechPaddlePaddle 的头像

    PaddlePaddle/PaddleSpeech

    12,626在 GitHub 上查看↗

    PaddleSpeech is a comprehensive toolkit of neural models for speech recognition, synthesis, and translation built on the PaddlePaddle deep learning framework. It provides a collection of frameworks and tools for converting spoken audio into written text, synthesizing natural audio from text, and performing direct speech translation. The toolkit includes specialized capabilities for keyword spotting to detect trigger words and speaker verification systems that extract unique voiceprints to identify and distinguish between individuals. It also features end-to-end translation tools that map audi

    Transforms generated spectral data into high-fidelity time-domain audio waveforms using neural vocoders.

    Pythonasrcode-switchconformer
    在 GitHub 上查看↗12,626
  • audiokit/audiokitaudiokit 的头像

    audiokit/AudioKit

    11,381在 GitHub 上查看↗

    AudioKit is an audio framework for iOS, macOS, and tvOS that provides tools for digital audio synthesis, signal processing, and audio analysis. It functions as a synthesis engine for generating audio waveforms and textures, a processing library for modifying tonal characteristics, and a toolkit for extracting frequency and amplitude data from sonic signals. The framework utilizes a modular node architecture and graph-based signal routing to connect audio generators, processors, and outputs. It wraps low-level audio primitives in high-level classes to facilitate sound generation and modificati

    Generates artificial audio waveforms and original sounds from scratch using digital synthesis.

    Swift
    在 GitHub 上查看↗11,381
  • aigc-audio/audiogptAIGC-Audio 的头像

    AIGC-Audio/AudioGPT

    10,174在 GitHub 上查看↗

    AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre

    Translates natural language descriptions into structured control signals to parameterize audio generation models.

    Pythonaudiogptmusic
    在 GitHub 上查看↗10,174
  • lmms/lmmsLMMS 的头像

    LMMS/lmms

    10,005在 GitHub 上查看↗

    LMMS is a digital audio workstation and MIDI sequencer designed for composing, arranging, and mixing music. It functions as a comprehensive production environment that integrates a MIDI sequencer, a sample-based synthesizer, and an audio mixing console. The project distinguishes itself through a versatile synthesis engine that includes additive synthesis, wavetable generation, and emulations of vintage hardware such as NES audio and FM chips. It also serves as a VST plugin host, allowing for the integration of third-party virtual instruments and audio effects via a standardized interface. Be

    Implements frequency modulation and additive synthesis to generate artificial audio signals and emulate hardware chips.

    C++dawhacktoberfestmidi
    在 GitHub 上查看↗10,005
  • espnet/espnetespnet 的头像

    espnet/espnet

    9,861在 GitHub 上查看↗

    ESPnet is a comprehensive speech processing toolkit and PyTorch-based trainer designed for building end-to-end speech recognition, synthesis, and translation models. It provides a structured framework for developing automatic speech recognition systems using transducer and encoder-decoder architectures, alongside engines for text-to-speech synthesis and speech translation pipelines. The project distinguishes itself through a recipe-based workflow execution system that ensures experimental reproducibility by running standardized sequences of scripts for data preparation and model training. It

    Generates melodic singing audio using non-autoregressive models with multi-speaker and multilingual support.

    Python
    在 GitHub 上查看↗9,861
上一个123下一个
  1. Home
  2. Graphics & Multimedia
  3. Media Processing and Analysis
  4. Audio Processing Systems
  5. Audio Synthesis

探索子标签

  • DAW Plugin InterfacesIntegrations that expose neural synthesis engines as plugins within digital audio workstations. **Distinct from Audio Synthesis:** Distinct from general audio synthesis: focuses on the plugin-based integration for DAW workflows.
  • Generative Composition SystemsSystems that generate structured musical sequences with long-term coherence using machine learning models. **Distinct from Audio Synthesis:** Distinct from general audio synthesis: focuses on the structural and compositional logic of musical sequences rather than raw signal generation.
  • MIDI-Driven Synthesis EnginesSynthesis systems that trigger and modulate audio output using standard MIDI performance data. **Distinct from Audio Synthesis:** Distinct from general audio synthesis: focuses on MIDI-based control and hardware integration.
  • MIDI-Driven Synthesis PlatformsPlatforms that trigger and control real-time generative audio synthesis using MIDI data. **Distinct from Audio Synthesis:** Distinct from general audio synthesis: focuses on the MIDI-driven control aspect of neural instruments.
  • Neural VocodersComputational models that transform generated spectral data into high-fidelity time-domain audio waveforms.
  • Parameter ModulationDynamic adjustment of synthesis model attributes using external control signals. **Distinct from Prompt-Driven Parameter Synthesis:** Focuses on real-time control voltage modulation of synthesis parameters rather than AI-driven prompt generation or granular slicing.
  • Reference-Driven Synthesis1 个子标签Audio generation that uses a specific reference sample to condition the output identity. **Distinct from Audio Synthesis:** Focuses on conditioning synthesis using a reference sample, whereas general audio synthesis covers all artificial signal generation.
  • Singing Voice Synthesis3 个子标签Specialized synthesis of melodic vocal performances based on text and timing prompts. **Distinct from Audio Synthesis:** Focuses specifically on melodic singing synthesis, distinct from general non-melodic audio synthesis.
  • Soundfont SynthesizersAudio synthesis systems that specifically use soundfont data to render realistic instrument sounds. **Distinct from Audio Synthesis:** Specializes in sample-based soundfont rendering rather than general neural or additive audio synthesis.
  • Timbre Morphing Tools2 个子标签Software for transforming the sonic characteristics of audio inputs to match target instrument profiles. **Distinct from Audio Synthesis:** Distinct from general audio synthesis: focuses specifically on real-time timbre transformation and style transfer.