1 repositorio
Loads a pretrained audio foundation model and runs inference on audio inputs to produce speech responses.
Distinct from Model Inference: Distinct from Model Inference: specifically targets audio foundation models for speech response generation, not general model inference.
Explore 1 awesome GitHub repository matching artificial intelligence & ml · Audio. Refine with filters or upvote what's useful.
Kimi-Audio is a large language model audio foundation model designed to understand audio input and generate high-fidelity speech responses in real time. It functions as a unified system encompassing a text-to-speech synthesis engine and a speech-to-text transcription tool. The project enables real-time audio conversations through a multi-modal conversation loop and chunk-wise streaming detokenization to reduce playback latency. It provides controls over speech speed, accent, and emotional tone during conversational audio generation. The system covers audio intelligence capabilities, includin
Loads a pretrained audio foundation model and runs inference on audio inputs to produce speech responses.