awesome-repositories.com
ब्लॉग
MCP
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टMCP सर्वरहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेस
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

14 रिपॉजिटरी

Awesome GitHub RepositoriesFull-Duplex Multimodal Interaction

Systems that process simultaneous visual, auditory, and text streams to enable fluid, real-time conversational interaction.

Distinct from Multimodal Processing: Distinct from Multimodal Processing: focuses on the full-duplex, real-time conversational aspect rather than just the integration of data modalities.

Explore 14 awesome GitHub repositories matching artificial intelligence & ml · Full-Duplex Multimodal Interaction. Refine with filters or upvote what's useful.

Awesome Full-Duplex Multimodal Interaction GitHub Repositories

AI के साथ बेहतरीन रिपॉजिटरी खोजें।हम AI का उपयोग करके सबसे सटीक रिपॉजिटरी खोजेंगे।
  • microsoft/ai-agents-for-beginnersmicrosoft का अवतार

    microsoft/ai-agents-for-beginners

    67,369GitHub पर देखें↗

    This project is a structured educational resource and technical guide for designing and implementing autonomous systems using large language models. It provides a comprehensive curriculum and code samples focused on agentic design patterns, autonomous development, and the creation of systems capable of planning and executing multi-step tasks. The resource details the implementation of agentic retrieval-augmented generation, where models autonomously plan and refine data searches. It covers a wide array of orchestrators and design patterns, including metacognitive reflection for self-correctin

    Supports processing and delivering information across text, voice, and sound for a consistent multimodal experience.

    Jupyter Notebookagentic-aiagentic-frameworkagentic-rag
    GitHub पर देखें↗67,369
  • sillytavern/sillytavernSillyTavern का अवतार

    SillyTavern/SillyTavern

    29,463GitHub पर देखें↗

    SillyTavern is a comprehensive interface and orchestration platform designed for immersive AI roleplay and interactive chat experiences. It functions as a unified gateway that connects users to a wide array of local and cloud-based large language models, providing a centralized environment to manage complex character personas, narrative context, and model-driven interactions. The platform distinguishes itself through its advanced prompt engineering and automation capabilities. It utilizes a sophisticated macro-based templating engine and vector-database retrieval to dynamically inject lore, c

    Integrates image generation, voice synthesis, and reactive character sprites for a rich multimodal experience.

    JavaScriptaichatllm
    GitHub पर देखें↗29,463
  • openbmb/minicpm-vOpenBMB का अवतार

    OpenBMB/MiniCPM-V

    25,653GitHub पर देखें↗

    MiniCPM-V is a multimodal large language model and vision-language system designed for complex visual and linguistic understanding. It functions as an on-device AI model, providing the capacity to process text, images, and video as a compact neural network. The project is specifically developed as an edge AI framework, utilizing quantization and weight sharding to run on memory-constrained mobile chipsets. This allows for the deployment of multimodal intelligence directly on mobile operating systems for local inference. Its capabilities cover multimodal content analysis of high-resolution im

    Processes simultaneous visual, auditory, and textual streams for fluid, full-duplex real-time conversations.

    Python
    GitHub पर देखें↗25,653
  • openbmb/minicpm-oOpenBMB का अवतार

    OpenBMB/MiniCPM-o

    23,850GitHub पर देखें↗

    MiniCPM-o is a multimodal large language model designed to function as a real-time conversational assistant on edge devices. By mapping text, image, video, and audio inputs into a unified latent space, the system enables simultaneous cross-modal reasoning and full-duplex interaction. It is built as an edge-side inference engine, utilizing quantized model weights to maintain high-performance processing on consumer hardware. The system distinguishes itself through its integrated speech synthesis and voice cloning capabilities, which allow for the generation of expressive, personalized vocal out

    Enables fluid, full-duplex interaction by processing simultaneous visual, auditory, and speech streams.

    Pythonminicpmminicpm-vmulti-modal
    GitHub पर देखें↗23,850
  • livekit/livekitlivekit का अवतार

    livekit/livekit

    19,358GitHub पर देखें↗

    LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with users through voice, video, and text. It provides a centralized, event-driven architecture to manage the entire lifecycle of automated participants, from initialization and session state management to graceful shutdown. By utilizing a selective forwarding unit, the platform efficiently routes media streams between participants and agents, ensuring low-latency communication and secure, token-based authentication for all connections. The platform distinguishes itself through it

    Coordinates simultaneous input and output across multiple modalities for natural communication.

    Gogolangmedia-serversfu
    GitHub पर देखें↗19,358
  • ten-framework/ten-frameworkTEN-framework का अवतार

    TEN-framework/ten-framework

    10,701GitHub पर देखें↗

    Ten Framework is a multimodal large language model agent framework designed for building low-latency conversational agents. It integrates voice, text, and visual inputs in real time to facilitate human interaction. The project includes a real-time speech processing pipeline for streaming transcription, voice activity detection, and speaker diarization. It also features an avatar synchronization engine that coordinates character lip animations and visual outputs with synthesized speech. The framework covers edge AI deployment through containerized packaging and direct integration with embedde

    Processes simultaneous audio and data flows to enable fluid, real-time conversational turns and natural interruptions.

    Pythonaimulti-modalreal-time
    GitHub पर देखें↗10,701
  • nvidia/personaplexNVIDIA का अवतार

    NVIDIA/personaplex

    10,030GitHub पर देखें↗

    Personaplex is an LLM speech-to-speech framework and conversational AI persona engine designed for real-time voice interfaces. It provides a system for defining AI identities and vocal characteristics through a combination of text-based role prompts and audio reference files. The project features a real-time AI voice interface that supports full-duplex human-AI dialogue, enabling multiple parties to speak and listen simultaneously via bidirectional audio streaming. It includes a GPU-accelerated audio processor and a speech-to-speech pipeline to facilitate low-latency conversations. The frame

    Enables real-time conversational systems where multiple parties can speak and listen simultaneously.

    Python
    GitHub पर देखें↗10,030
  • kyutai-labs/moshikyutai-labs का अवतार

    kyutai-labs/moshi

    9,672GitHub पर देखें↗

    Moshi is a real-time voice foundation model and speech-to-speech framework designed for bidirectional, low-latency conversations. It functions as a full-duplex voice interface that processes audio and text concurrently in a single stream, enabling natural human-machine dialogue without sequential processing delays. The system utilizes a neural audio codec to compress high-fidelity audio into low-bitrate tokens for efficient transmission. To manage complex responses and reasoning, it employs internal monologue modeling, which generates a hidden stream of thought tokens alongside audible speech

    Provides a full-duplex architecture that processes simultaneous audio and text streams for real-time conversation.

    Python
    GitHub पर देखें↗9,672
  • livekit/agentslivekit का अवतार

    livekit/agents

    9,379GitHub पर देखें↗

    This project is a framework for developing multimodal AI agents that function as programmable participants in real-time communication rooms. It enables the construction of agents that can see, hear, and speak by integrating speech-to-text, large language models, and text-to-speech pipelines to facilitate low-latency, natural conversations. The system is distinguished by its advanced orchestration of real-time media and conversational flow, including support for full-duplex speech, preemptive response generation, and sophisticated interruption management. It further differentiates itself throu

    Processes audio input and output simultaneously to support fluid, natural human-like conversations.

    Pythonagentsaiopenai
    GitHub पर देखें↗9,379
  • xinnan-tech/xiaozhi-esp32-serverxinnan-tech का अवतार

    xinnan-tech/xiaozhi-esp32-server

    8,627GitHub पर देखें↗

    This project is an AI voice assistant backend and gateway server designed to connect ESP32 hardware to large language models. It enables real-time conversational AI by processing streaming speech-to-text and text-to-speech interactions, allowing hardware devices to engage in natural language dialogue. The system is distinguished by a modular plugin framework that loads custom feature extensions at runtime and a retrieval-augmented generation engine that queries external knowledge bases for factual accuracy. It further personalizes interactions by using voiceprint mapping to identify individua

    Maintains persistent bidirectional connections for real-time streaming of audio and control data.

    JavaScriptdifyesp32mcp-server
    GitHub पर देखें↗8,627
  • openai/openai-realtime-agentsopenai का अवतार

    openai/openai-realtime-agents

    6,911GitHub पर देखें↗

    This project is a framework for building voice and text agents using the OpenAI Realtime API. It implements architectural patterns for multi-agent orchestration, hybrid model distribution, state-managed prompting, and real-time response validation. The framework utilizes a hybrid task distributor to split workloads between fast conversational models and high-intelligence models for complex reasoning. It employs an orchestration system that routes user requests between specialized agents using a graph to manage complex task requirements. Additional capabilities include a state machine prompt

    Maintains a persistent open connection for simultaneous audio and text streaming between client and server.

    TypeScript
    GitHub पर देखें↗6,911
  • modelengine-group/nexentModelEngine-Group का अवतार

    ModelEngine-Group/nexent

    5,265GitHub पर देखें↗

    Nexent is an enterprise AI control plane and LLM agent orchestration platform. It provides a zero-code environment for designing, deploying, and managing production AI agents through a multi-agent collaboration framework that coordinates specialized autonomous agents using standardized messaging protocols. The platform integrates the Model Context Protocol to connect agents with external tools, plugins, and services via a universal communication interface. It further distinguishes itself with a dedicated RAG knowledge base manager that imports unstructured documents and utilizes hybrid search

    Provides real-time conversational interaction processing across voice, text, images, and files.

    Pythonagentagentic-aiagentic-framework
    GitHub पर देखें↗5,265
  • qwenlm/qwen2.5-omniQwenLM का अवतार

    QwenLM/Qwen2.5-Omni

    4,026GitHub पर देखें↗

    Qwen2.5-Omni is an omnichannel multimodal large language model designed to process and generate content across text, audio, vision, and video. It functions as a real-time speech AI, utilizing an end-to-end architecture to maintain synchronous voice conversations with low-latency responses. The project emphasizes efficiency through quantized edge models, allowing for local inference on mobile hardware and resource-constrained devices. It employs 4-bit weight quantization, CPU-based process offloading, and on-demand weight loading to reduce GPU memory requirements. The system integrates specia

    Enables fluid, real-time conversational interactions using simultaneous visual, auditory, and text streams.

    Jupyter Notebook
    GitHub पर देखें↗4,026
  • soniqo/speech-swiftsoniqo का अवतार

    soniqo/speech-swift

    896GitHub पर देखें↗

    यह प्रोजेक्ट Apple Silicon के लिए विशेष रूप से इंजीनियर, ऑन-डिवाइस स्पीच रिकग्निशन, सिंथेसिस और ऑडियो प्रोसेसिंग के लिए एक व्यापक टूलकिट है। यह वास्तविक समय, फुल-डुप्लेक्स वॉयस एजेंट बनाने के लिए एक फ्रेमवर्क प्रदान करता है जो पूरी तरह से ऑफ़लाइन काम करते हैं, प्रदर्शन और गोपनीयता बनाए रखने के लिए नेटिव हार्डवेयर त्वरण का लाभ उठाते हैं। ऑप्टिमाइज़्ड मशीन लर्निंग मॉडल का उपयोग करके, लाइब्रेरी बाहरी क्लाउड सेवाओं पर निर्भरता के बिना जटिल ऑडियो कार्यों के स्थानीय निष्पादन को सक्षम बनाती है। यह लाइब्रेरी स्थानीय, उच्च-प्रदर्शन वॉयस इंटरैक्शन पर अपने विशेष फ़ोकस के माध्यम से खुद को अलग बनाती है। इसमें स्ट्रीमिंग ऑडियो पाइपलाइनों के लिए परिष्कृत ऑर्केस्ट्रेशन शामिल है, जो कम विलंबता के साथ वास्तविक समय ट्रांसक्रिप्शन, स्पीच सिंथेसिस और वॉयस क्लोनिंग की अनुमति देता है। सिस्टम को निरंतर, इंटरैक्टिव बातचीत को संभालने के लिए डिज़ाइन किया गया है, जिसमें ऑडियो फीडबैक लूप को रोकने और पर्सिस्टेंट स्ट्रीमिंग सत्रों को प्रबंधित करने के लिए इन-बिल्ट तंत्र शामिल हैं। मुख्य इंटरैक्शन से परे, प्रोजेक्ट ऑडियो एन्हांसमेंट और प्रबंधन क्षमताओं का एक विस्तृत सूट प्रदान करता है। यह स्पीकर डायराइजेशन और एम्बेडिंग निष्कर्षण के साथ-साथ सोर्स सेपरेशन, नॉइज़ रिडक्शन और ऑडियो अपसैंपलिंग सहित उन्नत सिग्नल प्रोसेसिंग का समर्थन करता है। फ्रेमवर्क व्यापक मॉडल प्रबंधन उपयोगिताएँ भी प्रदान करता है, जैसे क्वांटाइजेशन नियंत्रण, मेमोरी प्रबंधन और कस्टम मॉडल वेट लोडिंग के लिए समर्थन, यह सुनिश्चित करते हुए कि डेवलपर्स स्थानीय हार्डवेयर पर प्रोसेसिंग गति और संसाधन खपत को संतुलित कर सकें। प्रोजेक्ट में ऑडियो कार्यों को निष्पादित करने और मॉडल वेट्स को ऑप्टिमाइज़्ड प्रारूपों में बदलने के लिए एक कमांड-लाइन इंटरफ़ेस शामिल है। यह मानक उद्योग इंटरफ़ेस के साथ एकीकरण की सुविधा के लिए HTTP और WebSocket एंडपॉइंट्स को भी उजागर करता है।

    Processes continuous audio input to detect speech and generate spoken responses while allowing for real-time interruption.

    Swiftapple-siliconasrcoreml
    GitHub पर देखें↗896
  1. Home
  2. Artificial Intelligence & ML
  3. Multimodal Processing
  4. Full-Duplex Multimodal Interaction

सब-टैग एक्सप्लोर करें

  • Full-Duplex Voice InteractionsSystems that process continuous audio input to detect speech and generate spoken responses with real-time interruption. **Distinct from Full-Duplex Multimodal Interaction:** Distinct from Full-Duplex Multimodal Interaction: focuses specifically on voice-only conversational agents rather than multimodal streams.