awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

14 Repos

Awesome GitHub RepositoriesFull-Duplex Multimodal Interaction

Systems that process simultaneous visual, auditory, and text streams to enable fluid, real-time conversational interaction.

Distinct from Multimodal Processing: Distinct from Multimodal Processing: focuses on the full-duplex, real-time conversational aspect rather than just the integration of data modalities.

Explore 14 awesome GitHub repositories matching artificial intelligence & ml · Full-Duplex Multimodal Interaction. Refine with filters or upvote what's useful.

Awesome Full-Duplex Multimodal Interaction GitHub Repositories

Finde die besten Repos mit KI.Wir suchen mit KI nach den am besten passenden Repositories.
  • microsoft/ai-agents-for-beginnersAvatar von microsoft

    microsoft/ai-agents-for-beginners

    67,369Auf GitHub ansehen↗

    This project is a structured educational resource and technical guide for designing and implementing autonomous systems using large language models. It provides a comprehensive curriculum and code samples focused on agentic design patterns, autonomous development, and the creation of systems capable of planning and executing multi-step tasks. The resource details the implementation of agentic retrieval-augmented generation, where models autonomously plan and refine data searches. It covers a wide array of orchestrators and design patterns, including metacognitive reflection for self-correctin

    Supports processing and delivering information across text, voice, and sound for a consistent multimodal experience.

    Jupyter Notebookagentic-aiagentic-frameworkagentic-rag
    Auf GitHub ansehen↗67,369
  • sillytavern/sillytavernAvatar von SillyTavern

    SillyTavern/SillyTavern

    29,463Auf GitHub ansehen↗

    SillyTavern is a comprehensive interface and orchestration platform designed for immersive AI roleplay and interactive chat experiences. It functions as a unified gateway that connects users to a wide array of local and cloud-based large language models, providing a centralized environment to manage complex character personas, narrative context, and model-driven interactions. The platform distinguishes itself through its advanced prompt engineering and automation capabilities. It utilizes a sophisticated macro-based templating engine and vector-database retrieval to dynamically inject lore, c

    Integrates image generation, voice synthesis, and reactive character sprites for a rich multimodal experience.

    JavaScriptaichatllm
    Auf GitHub ansehen↗29,463
  • openbmb/minicpm-vAvatar von OpenBMB

    OpenBMB/MiniCPM-V

    25,653Auf GitHub ansehen↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Processes simultaneous visual, auditory, and textual streams for fluid, full-duplex real-time conversations.

    Python
    Auf GitHub ansehen↗25,653
  • openbmb/minicpm-oAvatar von OpenBMB

    OpenBMB/MiniCPM-o

    23,850Auf GitHub ansehen↗

    MiniCPM-o is a multimodal large language model designed to function as a real-time conversational assistant on edge devices. By mapping text, image, video, and audio inputs into a unified latent space, the system enables simultaneous cross-modal reasoning and full-duplex interaction. It is built as an edge-side inference engine, utilizing quantized model weights to maintain high-performance processing on consumer hardware. The system distinguishes itself through its integrated speech synthesis and voice cloning capabilities, which allow for the generation of expressive, personalized vocal out

    Enables fluid, full-duplex interaction by processing simultaneous visual, auditory, and speech streams.

    Pythonminicpmminicpm-vmulti-modal
    Auf GitHub ansehen↗23,850
  • livekit/livekitAvatar von livekit

    livekit/livekit

    19,358Auf GitHub ansehen↗

    LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with users through voice, video, and text. It provides a centralized, event-driven architecture to manage the entire lifecycle of automated participants, from initialization and session state management to graceful shutdown. By utilizing a selective forwarding unit, the platform efficiently routes media streams between participants and agents, ensuring low-latency communication and secure, token-based authentication for all connections. The platform distinguishes itself through it

    Coordinates simultaneous input and output across multiple modalities for natural communication.

    Gogolangmedia-serversfu
    Auf GitHub ansehen↗19,358
  • ten-framework/ten-frameworkAvatar von TEN-framework

    TEN-framework/ten-framework

    10,701Auf GitHub ansehen↗

    Ten Framework is a multimodal large language model agent framework designed for building low-latency conversational agents. It integrates voice, text, and visual inputs in real time to facilitate human interaction. The project includes a real-time speech processing pipeline for streaming transcription, voice activity detection, and speaker diarization. It also features an avatar synchronization engine that coordinates character lip animations and visual outputs with synthesized speech. The framework covers edge AI deployment through containerized packaging and direct integration with embedde

    Processes simultaneous audio and data flows to enable fluid, real-time conversational turns and natural interruptions.

    Pythonaimulti-modalreal-time
    Auf GitHub ansehen↗10,701
  • nvidia/personaplexAvatar von NVIDIA

    NVIDIA/personaplex

    10,030Auf GitHub ansehen↗

    Personaplex is an LLM speech-to-speech framework and conversational AI persona engine designed for real-time voice interfaces. It provides a system for defining AI identities and vocal characteristics through a combination of text-based role prompts and audio reference files. The project features a real-time AI voice interface that supports full-duplex human-AI dialogue, enabling multiple parties to speak and listen simultaneously via bidirectional audio streaming. It includes a GPU-accelerated audio processor and a speech-to-speech pipeline to facilitate low-latency conversations. The frame

    Enables real-time conversational systems where multiple parties can speak and listen simultaneously.

    Python
    Auf GitHub ansehen↗10,030
  • kyutai-labs/moshiAvatar von kyutai-labs

    kyutai-labs/moshi

    9,672Auf GitHub ansehen↗

    Moshi is a real-time voice foundation model and speech-to-speech framework designed for bidirectional, low-latency conversations. It functions as a full-duplex voice interface that processes audio and text concurrently in a single stream, enabling natural human-machine dialogue without sequential processing delays. The system utilizes a neural audio codec to compress high-fidelity audio into low-bitrate tokens for efficient transmission. To manage complex responses and reasoning, it employs internal monologue modeling, which generates a hidden stream of thought tokens alongside audible speech

    Provides a full-duplex architecture that processes simultaneous audio and text streams for real-time conversation.

    Python
    Auf GitHub ansehen↗9,672
  • livekit/agentsAvatar von livekit

    livekit/agents

    9,379Auf GitHub ansehen↗

    This project is a framework for developing multimodal AI agents that function as programmable participants in real-time communication rooms. It enables the construction of agents that can see, hear, and speak by integrating speech-to-text, large language models, and text-to-speech pipelines to facilitate low-latency, natural conversations. The system is distinguished by its advanced orchestration of real-time media and conversational flow, including support for full-duplex speech, preemptive response generation, and sophisticated interruption management. It further differentiates itself throu

    Processes audio input and output simultaneously to support fluid, natural human-like conversations.

    Pythonagentsaiopenai
    Auf GitHub ansehen↗9,379
  • xinnan-tech/xiaozhi-esp32-serverAvatar von xinnan-tech

    xinnan-tech/xiaozhi-esp32-server

    8,627Auf GitHub ansehen↗

    This project is an AI voice assistant backend and gateway server designed to connect ESP32 hardware to large language models. It enables real-time conversational AI by processing streaming speech-to-text and text-to-speech interactions, allowing hardware devices to engage in natural language dialogue. The system is distinguished by a modular plugin framework that loads custom feature extensions at runtime and a retrieval-augmented generation engine that queries external knowledge bases for factual accuracy. It further personalizes interactions by using voiceprint mapping to identify individua

    Maintains persistent bidirectional connections for real-time streaming of audio and control data.

    JavaScriptdifyesp32mcp-server
    Auf GitHub ansehen↗8,627
  • openai/openai-realtime-agentsAvatar von openai

    openai/openai-realtime-agents

    6,911Auf GitHub ansehen↗

    This project is a framework for building voice and text agents using the OpenAI Realtime API. It implements architectural patterns for multi-agent orchestration, hybrid model distribution, state-managed prompting, and real-time response validation. The framework utilizes a hybrid task distributor to split workloads between fast conversational models and high-intelligence models for complex reasoning. It employs an orchestration system that routes user requests between specialized agents using a graph to manage complex task requirements. Additional capabilities include a state machine prompt

    Maintains a persistent open connection for simultaneous audio and text streaming between client and server.

    TypeScript
    Auf GitHub ansehen↗6,911
  • modelengine-group/nexentAvatar von ModelEngine-Group

    ModelEngine-Group/nexent

    5,265Auf GitHub ansehen↗

    Nexent ist eine Enterprise-KI-Control-Plane und eine Plattform zur Orchestrierung von LLM-Agenten. Sie bietet eine Zero-Code-Umgebung zum Entwerfen, Bereitstellen und Verwalten von KI-Agenten in der Produktion durch ein Multi-Agent-Collaboration-Framework, das spezialisierte autonome Agenten mithilfe standardisierter Messaging-Protokolle koordiniert. Die Plattform integriert das Model Context Protocol, um Agenten mit externen Tools, Plugins und Diensten über eine universelle Kommunikationsschnittstelle zu verbinden. Sie zeichnet sich zudem durch einen dedizierten RAG-Knowledge-Base-Manager aus, der unstrukturierte Dokumente importiert und hybride Suche nutzt, um fundierten Kontext für Modellantworten bereitzustellen. Das System deckt ein breites Spektrum an Funktionen ab, darunter mandantenfähige rollenbasierte Zugriffskontrolle, multimodale Interaktion über Text, Sprache und Bilder sowie hybrides Vektor-Retrieval. Es enthält zudem einen Marktplatz für die Verteilung und Entdeckung von Agenten sowie Observability-Tools zur Erfassung von Ausführungs-Traces. Die Plattform unterstützt sichere Bereitstellung durch containerisierte Offline-Paketierung für Air-Gapped-Infrastrukturen.

    Provides real-time conversational interaction processing across voice, text, images, and files.

    Pythonagentagentic-aiagentic-framework
    Auf GitHub ansehen↗5,265
  • qwenlm/qwen2.5-omniAvatar von QwenLM

    QwenLM/Qwen2.5-Omni

    4,026Auf GitHub ansehen↗

    Qwen2.5-Omni ist ein multimodales Large Language Model für Omnichannel-Anwendungen, das Inhalte über Text, Audio, Bild und Video verarbeiten und generieren kann. Es fungiert als Echtzeit-Sprach-KI und nutzt eine End-to-End-Architektur, um synchrone Sprachkonversationen mit geringer Latenz zu ermöglichen. Das Projekt betont Effizienz durch quantisierte Edge-Modelle, die eine lokale Inferenz auf mobiler Hardware und ressourcenbeschränkten Geräten ermöglichen. Es verwendet 4-Bit-Gewichtungsquantisierung, CPU-basiertes Process-Offloading und On-Demand-Gewichtungsladung, um den GPU-Speicherbedarf zu senken. Das System integriert spezialisierte Encoder zur Analyse multimodaler Datenströme und verfügt über einen Streaming-Decoder für die Echtzeit-Sprachgenerierung. Es enthält zudem Funktionen zur Anpassung der Sprachausgabe, um die tonalen und geschlechtsspezifischen Eigenschaften des Audiosignals zu modifizieren.

    Enables fluid, real-time conversational interactions using simultaneous visual, auditory, and text streams.

    Jupyter Notebook
    Auf GitHub ansehen↗4,026
  • soniqo/speech-swiftAvatar von soniqo

    soniqo/speech-swift

    896Auf GitHub ansehen↗

    Dieses Projekt ist ein umfassendes Toolkit für On-Device-Spracherkennung, -Synthese und Audioverarbeitung, das speziell für Apple Silicon entwickelt wurde. Es bietet ein Framework für den Aufbau von Echtzeit-Voice-Agents mit Vollduplex-Funktionalität, die vollständig offline arbeiten und native Hardwarebeschleunigung nutzen, um Performance und Datenschutz zu wahren. Durch den Einsatz optimierter Machine-Learning-Modelle ermöglicht die Bibliothek die lokale Ausführung komplexer Audioaufgaben ohne Abhängigkeit von externen Cloud-Diensten. Die Bibliothek zeichnet sich durch ihren spezialisierten Fokus auf lokale, hochperformante Sprachinteraktion aus. Sie enthält eine ausgefeilte Orchestrierung für Streaming-Audio-Pipelines, die Echtzeit-Transkription, Sprachsynthese und Voice-Cloning mit geringer Latenz ermöglicht. Das System ist für die Handhabung kontinuierlicher, interaktiver Konversationen konzipiert und verfügt über integrierte Mechanismen zur Vermeidung von Audio-Feedback-Schleifen und zur Verwaltung persistenter Streaming-Sitzungen. Über die Kerninteraktion hinaus bietet das Projekt eine breite Palette an Audio-Enhancement- und Management-Funktionen. Es unterstützt fortgeschrittene Signalverarbeitung, einschließlich Quellentrennung, Rauschunterdrückung und Audio-Upsampling, neben Tools für Sprecher-Diarisierung und Embedding-Extraktion. Das Framework bietet zudem umfangreiche Modellmanagement-Utilities, wie z. B. Quantisierungskontrollen, Speicherverwaltung und Unterstützung für das Laden benutzerdefinierter Modellgewichte, um sicherzustellen, dass Entwickler Verarbeitungsgeschwindigkeit und Ressourcenverbrauch auf lokaler Hardware ausbalancieren können. Das Projekt enthält eine CLI für die Ausführung von Audioaufgaben und die Konvertierung von Modellgewichten in optimierte Formate. Es stellt zudem HTTP- und WebSocket-Endpunkte bereit, um die Integration mit Standard-Industrieschnittstellen zu erleichtern.

    Processes continuous audio input to detect speech and generate spoken responses while allowing for real-time interruption.

    Swiftapple-siliconasrcoreml
    Auf GitHub ansehen↗896
  1. Home
  2. Artificial Intelligence & ML
  3. Multimodal Processing
  4. Full-Duplex Multimodal Interaction

Unter-Tags erkunden

  • Full-Duplex Voice InteractionsSystems that process continuous audio input to detect speech and generate spoken responses with real-time interruption. **Distinct from Full-Duplex Multimodal Interaction:** Distinct from Full-Duplex Multimodal Interaction: focuses specifically on voice-only conversational agents rather than multimodal streams.