14 Repos
Systems that process simultaneous visual, auditory, and text streams to enable fluid, real-time conversational interaction.
Distinct from Multimodal Processing: Distinct from Multimodal Processing: focuses on the full-duplex, real-time conversational aspect rather than just the integration of data modalities.
Explore 14 awesome GitHub repositories matching artificial intelligence & ml · Full-Duplex Multimodal Interaction. Refine with filters or upvote what's useful.
This project is a structured educational resource and technical guide for designing and implementing autonomous systems using large language models. It provides a comprehensive curriculum and code samples focused on agentic design patterns, autonomous development, and the creation of systems capable of planning and executing multi-step tasks. The resource details the implementation of agentic retrieval-augmented generation, where models autonomously plan and refine data searches. It covers a wide array of orchestrators and design patterns, including metacognitive reflection for self-correctin
Supports processing and delivering information across text, voice, and sound for a consistent multimodal experience.
SillyTavern is a comprehensive interface and orchestration platform designed for immersive AI roleplay and interactive chat experiences. It functions as a unified gateway that connects users to a wide array of local and cloud-based large language models, providing a centralized environment to manage complex character personas, narrative context, and model-driven interactions. The platform distinguishes itself through its advanced prompt engineering and automation capabilities. It utilizes a sophisticated macro-based templating engine and vector-database retrieval to dynamically inject lore, c
Integrates image generation, voice synthesis, and reactive character sprites for a rich multimodal experience.
MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im
Processes simultaneous visual, auditory, and textual streams for fluid, full-duplex real-time conversations.
MiniCPM-o is a multimodal large language model designed to function as a real-time conversational assistant on edge devices. By mapping text, image, video, and audio inputs into a unified latent space, the system enables simultaneous cross-modal reasoning and full-duplex interaction. It is built as an edge-side inference engine, utilizing quantized model weights to maintain high-performance processing on consumer hardware. The system distinguishes itself through its integrated speech synthesis and voice cloning capabilities, which allow for the generation of expressive, personalized vocal out
Enables fluid, full-duplex interaction by processing simultaneous visual, auditory, and speech streams.
LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with users through voice, video, and text. It provides a centralized, event-driven architecture to manage the entire lifecycle of automated participants, from initialization and session state management to graceful shutdown. By utilizing a selective forwarding unit, the platform efficiently routes media streams between participants and agents, ensuring low-latency communication and secure, token-based authentication for all connections. The platform distinguishes itself through it
Coordinates simultaneous input and output across multiple modalities for natural communication.
Ten Framework is a multimodal large language model agent framework designed for building low-latency conversational agents. It integrates voice, text, and visual inputs in real time to facilitate human interaction. The project includes a real-time speech processing pipeline for streaming transcription, voice activity detection, and speaker diarization. It also features an avatar synchronization engine that coordinates character lip animations and visual outputs with synthesized speech. The framework covers edge AI deployment through containerized packaging and direct integration with embedde
Processes simultaneous audio and data flows to enable fluid, real-time conversational turns and natural interruptions.
Personaplex is an LLM speech-to-speech framework and conversational AI persona engine designed for real-time voice interfaces. It provides a system for defining AI identities and vocal characteristics through a combination of text-based role prompts and audio reference files. The project features a real-time AI voice interface that supports full-duplex human-AI dialogue, enabling multiple parties to speak and listen simultaneously via bidirectional audio streaming. It includes a GPU-accelerated audio processor and a speech-to-speech pipeline to facilitate low-latency conversations. The frame
Enables real-time conversational systems where multiple parties can speak and listen simultaneously.
Moshi is a real-time voice foundation model and speech-to-speech framework designed for bidirectional, low-latency conversations. It functions as a full-duplex voice interface that processes audio and text concurrently in a single stream, enabling natural human-machine dialogue without sequential processing delays. The system utilizes a neural audio codec to compress high-fidelity audio into low-bitrate tokens for efficient transmission. To manage complex responses and reasoning, it employs internal monologue modeling, which generates a hidden stream of thought tokens alongside audible speech
Provides a full-duplex architecture that processes simultaneous audio and text streams for real-time conversation.
This project is a framework for developing multimodal AI agents that function as programmable participants in real-time communication rooms. It enables the construction of agents that can see, hear, and speak by integrating speech-to-text, large language models, and text-to-speech pipelines to facilitate low-latency, natural conversations. The system is distinguished by its advanced orchestration of real-time media and conversational flow, including support for full-duplex speech, preemptive response generation, and sophisticated interruption management. It further differentiates itself throu
Processes audio input and output simultaneously to support fluid, natural human-like conversations.
This project is an AI voice assistant backend and gateway server designed to connect ESP32 hardware to large language models. It enables real-time conversational AI by processing streaming speech-to-text and text-to-speech interactions, allowing hardware devices to engage in natural language dialogue. The system is distinguished by a modular plugin framework that loads custom feature extensions at runtime and a retrieval-augmented generation engine that queries external knowledge bases for factual accuracy. It further personalizes interactions by using voiceprint mapping to identify individua
Maintains persistent bidirectional connections for real-time streaming of audio and control data.
This project is a framework for building voice and text agents using the OpenAI Realtime API. It implements architectural patterns for multi-agent orchestration, hybrid model distribution, state-managed prompting, and real-time response validation. The framework utilizes a hybrid task distributor to split workloads between fast conversational models and high-intelligence models for complex reasoning. It employs an orchestration system that routes user requests between specialized agents using a graph to manage complex task requirements. Additional capabilities include a state machine prompt
Maintains a persistent open connection for simultaneous audio and text streaming between client and server.
Nexent ist eine Enterprise-KI-Control-Plane und eine Plattform zur Orchestrierung von LLM-Agenten. Sie bietet eine Zero-Code-Umgebung zum Entwerfen, Bereitstellen und Verwalten von KI-Agenten in der Produktion durch ein Multi-Agent-Collaboration-Framework, das spezialisierte autonome Agenten mithilfe standardisierter Messaging-Protokolle koordiniert. Die Plattform integriert das Model Context Protocol, um Agenten mit externen Tools, Plugins und Diensten über eine universelle Kommunikationsschnittstelle zu verbinden. Sie zeichnet sich zudem durch einen dedizierten RAG-Knowledge-Base-Manager aus, der unstrukturierte Dokumente importiert und hybride Suche nutzt, um fundierten Kontext für Modellantworten bereitzustellen. Das System deckt ein breites Spektrum an Funktionen ab, darunter mandantenfähige rollenbasierte Zugriffskontrolle, multimodale Interaktion über Text, Sprache und Bilder sowie hybrides Vektor-Retrieval. Es enthält zudem einen Marktplatz für die Verteilung und Entdeckung von Agenten sowie Observability-Tools zur Erfassung von Ausführungs-Traces. Die Plattform unterstützt sichere Bereitstellung durch containerisierte Offline-Paketierung für Air-Gapped-Infrastrukturen.
Provides real-time conversational interaction processing across voice, text, images, and files.
Qwen2.5-Omni ist ein multimodales Large Language Model für Omnichannel-Anwendungen, das Inhalte über Text, Audio, Bild und Video verarbeiten und generieren kann. Es fungiert als Echtzeit-Sprach-KI und nutzt eine End-to-End-Architektur, um synchrone Sprachkonversationen mit geringer Latenz zu ermöglichen. Das Projekt betont Effizienz durch quantisierte Edge-Modelle, die eine lokale Inferenz auf mobiler Hardware und ressourcenbeschränkten Geräten ermöglichen. Es verwendet 4-Bit-Gewichtungsquantisierung, CPU-basiertes Process-Offloading und On-Demand-Gewichtungsladung, um den GPU-Speicherbedarf zu senken. Das System integriert spezialisierte Encoder zur Analyse multimodaler Datenströme und verfügt über einen Streaming-Decoder für die Echtzeit-Sprachgenerierung. Es enthält zudem Funktionen zur Anpassung der Sprachausgabe, um die tonalen und geschlechtsspezifischen Eigenschaften des Audiosignals zu modifizieren.
Enables fluid, real-time conversational interactions using simultaneous visual, auditory, and text streams.
Dieses Projekt ist ein umfassendes Toolkit für On-Device-Spracherkennung, -Synthese und Audioverarbeitung, das speziell für Apple Silicon entwickelt wurde. Es bietet ein Framework für den Aufbau von Echtzeit-Voice-Agents mit Vollduplex-Funktionalität, die vollständig offline arbeiten und native Hardwarebeschleunigung nutzen, um Performance und Datenschutz zu wahren. Durch den Einsatz optimierter Machine-Learning-Modelle ermöglicht die Bibliothek die lokale Ausführung komplexer Audioaufgaben ohne Abhängigkeit von externen Cloud-Diensten. Die Bibliothek zeichnet sich durch ihren spezialisierten Fokus auf lokale, hochperformante Sprachinteraktion aus. Sie enthält eine ausgefeilte Orchestrierung für Streaming-Audio-Pipelines, die Echtzeit-Transkription, Sprachsynthese und Voice-Cloning mit geringer Latenz ermöglicht. Das System ist für die Handhabung kontinuierlicher, interaktiver Konversationen konzipiert und verfügt über integrierte Mechanismen zur Vermeidung von Audio-Feedback-Schleifen und zur Verwaltung persistenter Streaming-Sitzungen. Über die Kerninteraktion hinaus bietet das Projekt eine breite Palette an Audio-Enhancement- und Management-Funktionen. Es unterstützt fortgeschrittene Signalverarbeitung, einschließlich Quellentrennung, Rauschunterdrückung und Audio-Upsampling, neben Tools für Sprecher-Diarisierung und Embedding-Extraktion. Das Framework bietet zudem umfangreiche Modellmanagement-Utilities, wie z. B. Quantisierungskontrollen, Speicherverwaltung und Unterstützung für das Laden benutzerdefinierter Modellgewichte, um sicherzustellen, dass Entwickler Verarbeitungsgeschwindigkeit und Ressourcenverbrauch auf lokaler Hardware ausbalancieren können. Das Projekt enthält eine CLI für die Ausführung von Audioaufgaben und die Konvertierung von Modellgewichten in optimierte Formate. Es stellt zudem HTTP- und WebSocket-Endpunkte bereit, um die Integration mit Standard-Industrieschnittstellen zu erleichtern.
Processes continuous audio input to detect speech and generate spoken responses while allowing for real-time interruption.