1 repository
Systems that allow providing text-based input as a substitute for other modalities like speech.
Distinguishing note: None of the candidates cover the logic of using text as a modality alternative for conversational AI agents.
Explore 1 awesome GitHub repository matching artificial intelligence & ml · Multimodal Input Alternatives. Refine with filters or upvote what's useful.
Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech systems. It utilizes a frame-based data pipeline to route audio, video, and text through a modular sequence of processors, enabling the orchestration of low-latency conversational AI. The project is distinguished by its ability to coordinate complex multimodal services, including speech-to-text, language models, and text-to-speech, within a single pipeline. It features semantic voice activity detection for natural turn-taking, state-machine conversation flows for dialogue manag
Allows sending text content to the bot as an alternative to audio input.