1 repositorio
Maintains context across multiple spoken exchanges, generating both text and audio replies.
Distinct from Multi-Turn Agent Conversations: Distinct from Multi-Turn Agent Conversations: focuses on speech-based multi-turn conversations with audio input/output, not agent-to-agent dialogues.
Explore 1 awesome GitHub repository matching artificial intelligence & ml · Multi-Turn Speech Conversations. Refine with filters or upvote what's useful.
Kimi-Audio is a large language model audio foundation model designed to understand audio input and generate high-fidelity speech responses in real time. It functions as a unified system encompassing a text-to-speech synthesis engine and a speech-to-text transcription tool. The project enables real-time audio conversations through a multi-modal conversation loop and chunk-wise streaming detokenization to reduce playback latency. It provides controls over speech speed, accent, and emotional tone during conversational audio generation. The system covers audio intelligence capabilities, includin
Maintains context across multiple spoken exchanges, generating both text and audio replies.