awesome-repositories.com
Blog
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektÜber unsRanking-MethodikPresseMCP-Server
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
duixcom avatar

duixcom/Duix-Mobile

0
View on GitHub↗
8,093 Stars·1,196 Forks·C++·8 Aufrufewww.duix.com↗

Duix Mobile

Duix-Mobile is a software development kit for deploying real-time conversational AI characters on mobile devices. It enables the creation of interactive digital humans capable of fluid voice-to-voice interactions, featuring low-latency speech recognition and synchronized lip movements.

The project distinguishes itself through the ability to integrate custom external language models and speech providers to define an avatar's intelligence and voice. It supports the generation of real-time multilingual subtitles and provides mechanisms to track the training status of newly created digital characters.

The system covers a broad range of capabilities, including session lifecycle management, bidirectional media streaming, and audio-driven animation. It includes tools for automatic speech recognition, camera access control, and token-based session authentication. Visual rendering is handled via web-container and custom view options to decouple assets from the native host.

The SDK includes build-time optimizations for Android to prevent the obfuscation of critical system classes.

Features

  • Interactive Video Avatar Generators - Enables the deployment and management of real-time digital humans for voiced conversations on mobile devices.
  • Session Initializers - Bootstraps new interactive digital human sessions and establishes communication rooms for connectivity.
  • AI Model Integrations - Provides capabilities for connecting external language models and speech providers to customize avatar intelligence and voice.
  • AI Provider Integrations - Provides connectors for integrating external AI model providers to power the avatar's intelligence and voice.
  • Automatic Speech Recognition - Converts spoken user audio into written text in real time to drive interactions via callbacks.
  • Conversational Speech Capture - Records audio input or uses automatic speech recognition to convert user voice into text for AI interaction.
  • Interactive AI Conversations - Facilitates interactive Q&A by sending user questions to a digital human and receiving generated responses.
  • Interactive Session Launchers - Starts real-time sessions after asset verification and optionally enables automatic speech recognition.
  • Conversational Response Generation - Generates AI-driven responses to user questions using connected language models with optional memory support.
  • Real-Time Conversational AI Frameworks - Implements a framework for low-latency voice-to-voice interactions integrating speech recognition and synchronized animation.
  • Realtime AI Session Managers - Manages persistent, bidirectional communication sessions with AI models to support low-latency voice interactions.
  • Conversational Audio Streams - Implements real-time processing pipelines for voice-based interaction with support for simultaneous playback and interruptions.
  • Text To Speech - Commands digital characters to synthesize and speak text or audio files in real-time.
  • Lip-Sync Stream Synchronization - Synchronizes incoming audio and video data to align the digital human's lip movements with synthesized speech.
  • Real-Time Media Streaming - Implements low-latency bidirectional streaming of audio and video for real-time interaction with digital humans.
  • Digital Human Connection Management - Establishes real-time connections to digital human instances using unique conversation identifiers.
  • Session Lifecycle Management - Handles the complex asynchronous sequence of initializing, connecting, and terminating digital human sessions.
  • Real-time Audio Capture Protocols - Captures sound from the device microphone using real-time communication protocols to enable interactive voice sessions.
  • API Request Authentication - Verifies identity and authorizes access to AI services using secure tokens passed in request headers.
  • Conversational Session Managers - Implements comprehensive management of the lifecycle, connection, and synchronization for conversational AI sessions.
  • Audio-Driven Animation Engines - Implements systems that synchronize character movements and mouth shapes to match live audio rhythm in real time.
  • Realtime Avatar Renderers - Initializes real-time sessions and renders interactive avatars within a web container for user engagement.
  • Avatar Speech Control - Triggers digital humans to speak using text or audio URLs, including the ability to interrupt current playback.
  • Lip Synchronization Engines - Provides engines that align digital character mouth movements with synthesized audio for visual alignment.
  • Speech Interruption Management - Provides the capability to stop current audio output immediately to halt the digital human's speaking state.
  • Automated Subtitle Generators - Produces real-time multilingual text overlays to accompany the spoken responses of digital characters.
  • Avatar Lifecycle Events - Tracks loading progress, speech recognition results, and connection changes via registered callbacks.
  • Audio Recording - Captures audio input from device microphones to provide voice input for interactive sessions.
  • Media Streaming - Retrieves local and remote audio and video streams for custom rendering and media processing.
  • Multilingual Captioning - Supports digital characters that can interact in multiple languages with synchronized real-time subtitles.
  • Audio Modality Controls - Provides logic for dynamically enabling or disabling microphone input and avatar audio output streams.
  • Web-View Containers - Renders interactive digital humans within a web-view container to decouple visual assets from the native host.
  • Session Identifiers - Provides unique session identifiers to enable programmatic connection and interaction with digital human instances.
  • Bulk Communication Session Termination - Ends all current digital character interactions associated with an application to clear active connections.
  • Communication Session Termination - Provides mechanisms for ending real-time media sessions and cleaning up associated agent resources.
  • Targeted Avatar Session Termination - Ends an active interactive session using a unique identifier to stop the character's operation.
  • Token-Based Authentication - Uses secure header tokens to verify application identity and authorize real-time communication sessions.
  • Event Dispatchers - Uses an event dispatcher to notify the host application of session state changes and speech recognition results.
  • User Session Monitors - Retrieves a list of all active in-call sessions to track real-time avatar usage.
  • Concurrency Monitoring - Tracks the number of active concurrent sessions for applications and users to manage system load.
  • Session Rendering Configurations - Configures rendering containers and authentication to establish connections between the device and the interactive platform.
  • Avatar Visual Controls - Provides controls for starting visual playback with options for background removal and audio muting.

Star-Verlauf

Star-Verlauf für duixcom/duix-mobileStar-Verlauf für duixcom/duix-mobile

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Open-Source-Alternativen zu Duix Mobile

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit Duix Mobile.
  • livekit/agentsAvatar von livekit

    livekit/agents

    9,379Auf GitHub ansehen↗

    This project is a framework for developing multimodal AI agents that function as programmable participants in real-time communication rooms. It enables the construction of agents that can see, hear, and speak by integrating speech-to-text, large language models, and text-to-speech pipelines to facilitate low-latency, natural conversations. The system is distinguished by its advanced orchestration of real-time media and conversational flow, including support for full-duplex speech, preemptive response generation, and sophisticated interruption management. It further differentiates itself throu

    Pythonagentsaiopenai
    Auf GitHub ansehen↗9,379
  • lipku/livetalkingAvatar von lipku

    lipku/LiveTalking

    8,042Auf GitHub ansehen↗

    LiveTalking is an interactive talking head engine and AI avatar management platform designed to synchronize synthetic speech with facial movements. It functions as a real-time orchestrator that connects large language models and text-to-speech services to neural-rendered digital humans. The project distinguishes itself through low-latency streaming capabilities and the ability to handle real-time conversational interruptions. It supports advanced audio-visual customization, including human voice cloning and the ability to drive avatar expressions using real-time webcam data. The platform cov

    Pythonaigcdigihumandigital-human
    Auf GitHub ansehen↗8,042
  • pipecat-ai/pipecatAvatar von pipecat-ai

    pipecat-ai/pipecat

    12,846Auf GitHub ansehen↗

    Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech systems. It utilizes a frame-based data pipeline to route audio, video, and text through a modular sequence of processors, enabling the orchestration of low-latency conversational AI. The project is distinguished by its ability to coordinate complex multimodal services, including speech-to-text, language models, and text-to-speech, within a single pipeline. It features semantic voice activity detection for natural turn-taking, state-machine conversation flows for dialogue manag

    Pythonaichatbot-frameworkchatbots
    Auf GitHub ansehen↗12,846
  • livekit/livekitAvatar von livekit

    livekit/livekit

    19,358Auf GitHub ansehen↗

    LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with users through voice, video, and text. It provides a centralized, event-driven architecture to manage the entire lifecycle of automated participants, from initialization and session state management to graceful shutdown. By utilizing a selective forwarding unit, the platform efficiently routes media streams between participants and agents, ensuring low-latency communication and secure, token-based authentication for all connections. The platform distinguishes itself through it

    Gogolangmedia-serversfu
    Auf GitHub ansehen↗19,358
Alle 30 Alternativen zu Duix Mobile anzeigen→

Häufig gestellte Fragen

Was macht duixcom/duix-mobile?

Duix-Mobile is a software development kit for deploying real-time conversational AI characters on mobile devices. It enables the creation of interactive digital humans capable of fluid voice-to-voice interactions, featuring low-latency speech recognition and synchronized lip movements.

Was sind die Hauptfunktionen von duixcom/duix-mobile?

Die Hauptfunktionen von duixcom/duix-mobile sind: Interactive Video Avatar Generators, Session Initializers, AI Model Integrations, AI Provider Integrations, Automatic Speech Recognition, Conversational Speech Capture, Interactive AI Conversations, Interactive Session Launchers.

Welche Open-Source-Alternativen gibt es zu duixcom/duix-mobile?

Open-Source-Alternativen zu duixcom/duix-mobile sind unter anderem: livekit/agents — This project is a framework for developing multimodal AI agents that function as programmable participants in… lipku/livetalking — LiveTalking is an interactive talking head engine and AI avatar management platform designed to synchronize synthetic… pipecat-ai/pipecat — Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech… livekit/livekit — LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with… getstream/vision-agents. kilo-org/kilocode — Kilocode is an autonomous engineering platform designed to orchestrate AI agents for complex software development…