2 dépôts
Neural processes for predicting and restoring missing segments of audio recordings.
Distinct from Audio Processing: Specifically targets the restoration of missing audio segments, whereas general audio processing covers a broader range of manipulations.
Explore 2 awesome GitHub repositories matching graphics & multimedia · Audio Gap Infilling. Refine with filters or upvote what's useful.
AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre
Restores missing segments of sound recordings by predicting and inserting the absent audio data.
VoiceCraft is a neural speech generation and manipulation system consisting of a text-to-speech system, a voice cloning tool, and an audio inpainting engine. It uses a large language model approach to synthesize high-fidelity audio from text and replicate speaker identities. The system provides zero-shot voice cloning and speech editing capabilities, allowing users to modify spoken content within existing recordings. This includes an audio inpainting engine that replaces specific sections of audio with new speech while preserving the original acoustic characteristics and speaker identity. Th
Provides a neural engine for predicting and restoring missing audio segments to modify spoken content.