2 مستودعات
Identifying and tagging the current speaker in real-time multi-modal streams.
Distinct from Speaker Diarization: Distinct from Speaker Diarization: focuses on real-time 'active' status using multi-sensor input rather than segmenting a recorded audio file.
Explore 2 awesome GitHub repositories matching artificial intelligence & ml · Active Speaker Detection. Refine with filters or upvote what's useful.
jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput. The project distinguishes itself through high-performance streaming analytics and the ability to execute concurrent AI pipelines on auto-grade silicon. It provides specialized support for multi-sensor stream processing, utilizing zero-copy data transport to load camera frames directly into GPU memory. The codebase covers a broad surface of capabiliti
NVIDIA detects and tags multiple speakers in live broadcast workflows using multi-camera and multi-microphone inputs.
Detects and tags which person is speaking in real-time across multiple camera and microphone feeds.