1 个仓库
Audio frameworks that use large language models to decompose goals and select appropriate neural audio tools.
Distinct from Audio Processing Frameworks: Focuses on the LLM-based control plane for tool selection, rather than just the underlying processing infrastructure.
Explore 1 awesome GitHub repository matching graphics & multimedia · LLM-Driven Orchestration. Refine with filters or upvote what's useful.
AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre
Uses large language models to orchestrate neural audio pipelines for generation and processing tasks.