awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目关于排名机制媒体报道MCP 服务器
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
huggingface avatar

huggingface/speech-to-speech

0
View on GitHub↗
4,895 星标·584 分支·Python·Apache-2.0·9 次浏览

Speech To Speech

该项目是一个用于构建本地语音助手和实时音频流服务器的框架。它作为一个容器化推理引擎和多语言语音流水线,编排语音转文字 (STT)、语言模型和文字转语音 (TTS) 组件,将语音输入转换为语音输出。

该系统以其使用基于 WebSocket 的双向流来实现低延迟交互而著称。它具有一个语音活动检测系统,可管理语音边界并处理助手播放期间的用户打断。它还支持通过音频预设进行自定义语音克隆,以及交换模型检查点或外部 API 进行识别和合成的能力。

该框架涵盖了广泛的功能面,包括异步音频缓冲、事件驱动的轮次管理和基于 Schema 的工具执行。它支持多语言对话管理,并通过基于线程的流水线隔离运行并发会话。

该项目提供针对 x86 和 ARM64 架构优化的容器镜像。

Features

  • Voice Agents - Provides a framework for building local voice agents that convert spoken input into spoken output.
  • Voice Assistants - Provides a framework for building local voice assistants using open-source models for speech-to-speech interaction.
  • Turn Event Emitters - Emits events when user speech boundaries are detected to trigger the corresponding transcription and synthesis stages.
  • Voice Activity Detection - Uses silence and duration thresholds to automatically identify speech segments within audio streams.
  • Voice Pipelines - Coordinates the sequence of voice detection, transcription, language processing, and synthesis into a fluid loop.
  • Voice Interaction Management - Coordinates the entire voice interaction loop, from speech detection and recognition to response synthesis.
  • Model Provider Integrations - Connects to diverse language model providers through a unified interface supporting local and cloud-based APIs.
  • Real-Time Transcription - Provides instantaneous conversion of live audio streams into text transcripts displayed in real-time.
  • Speech Boundary Detection - Identifies exact start and end timestamps of human speech to trigger transcription and manage conversational turns.
  • Modular Pipeline Orchestrators - Provides a modular orchestrator that separates voice activity detection, transcription, and synthesis into independent processing components.
  • Real-Time Speech Processing - Coordinates low-latency workflows sequencing voice activity detection, transcription, language processing, and synthesis.
  • Speech-to-Text and Text-to-Speech Integrations - Provides a configurable pipeline for integrating speech-to-text, language models, and text-to-speech backends.
  • Voice Activity Detection - Identifies speech boundaries to manage conversational turn-taking and handle user barge-in interruptions.
  • Bidirectional WebSocket Streaming - Implements bidirectional WebSocket streaming for low-latency exchange of raw audio and text data between client and server.
  • Audio Streaming Servers - Implements a dedicated audio streaming server for bidirectional raw audio delivery with low-latency turn-taking.
  • Bidirectional Audio Transports - Implements WebSocket-based bidirectional audio transports for low-latency turn-taking and live transcription.
  • Voice Interaction Interfaces - Provides a conversational interface that manages turn-taking and user interruptions in real-time audio streams.
  • Containerized Deployments - Packages the voice agent environment and dependencies into portable container images for consistent deployment.
  • Threshold Tunings - Allows adjustment of speech duration and silence thresholds to balance responsiveness with accurate turn segmentation.
  • Model Checkpoint Swapping - Enables integrating different model checkpoints or external APIs to optimize for specific hardware or latency requirements.
  • Voice Cloning Tools - Generates synthetic speech that mimics specific speakers using custom audio files or voice presets.
  • Local Voice Control Interfaces - Executes full speech-to-speech pipelines on local hardware using accelerators to minimize interaction latency.
  • Local Inference Engines - Acts as a local inference engine for running speech-to-speech models on consumer hardware.
  • Speech-to-Speech Translation - Implements a multilingual pipeline that detects spoken languages and converts audio to audio across different tongues.
  • Tool Call Executions - Extracts structured tool calls from model output using JSON schemas to trigger external actions.
  • Multilingual Audio Processing - Features a multilingual audio processing chain that handles language detection and audio-to-audio conversion.
  • Multilingual Conversational Interaction - Supports automatic spoken language detection and manages transcription and synthesis across multiple languages.
  • Barge-In Handlers - Implements handlers that stop current speech output immediately when a user barge-in event is detected.
  • Tool-Calling Schemas - Implements schema-based tool calling to allow language models to trigger external functions and retrieve data during voice conversations.
  • Real-Time Voice Protocols - Exposes a WebSocket endpoint supporting live transcription and turn-taking based on standardized real-time specifications.
  • Audio Buffers - Uses circular memory buffers for low-latency capture and real-time processing of raw audio input.
  • Model Agnostic Interfaces - Decouples high-level API calls from specific model implementations, allowing local transformers or cloud APIs to be used interchangeably.

Star 历史

huggingface/speech-to-speech 的 Star 历史图表huggingface/speech-to-speech 的 Star 历史图表

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Start searching with AI

包含 Speech To Speech 的精选搜索

收录 Speech To Speech 的精选合集。
  • 实时语音智能体框架
  • 自托管自然语音合成引擎
  • 自托管实时会议转录工具

Speech To Speech 的开源替代方案

相似的开源项目,按与 Speech To Speech 的功能重合度排序。
  • getstream/vision-agentsGetStream 的头像

    GetStream/Vision-Agents

    6,029在 GitHub 上查看↗
    Pythonagentic-aiagentsai
    在 GitHub 上查看↗6,029
  • vocodedev/vocode-corevocodedev 的头像

    vocodedev/vocode-core

    3,693在 GitHub 上查看↗

    Vocode-core is a framework for building real-time conversational AI voice agents. It serves as a conversational orchestrator and pipeline that integrates speech-to-text, large language models, and text-to-speech services to enable low-latency voice interactions. The project features a provider-agnostic interface that allows for swappable speech and language model providers, including support for both cloud APIs and local binaries. It distinguishes itself through a specialized telephony integration layer that enables agents to be deployed across phone lines, WebRTC, and virtual meeting platfor

    Python
    在 GitHub 上查看↗3,693
  • elevenlabs/elevenlabs-pythonelevenlabs 的头像

    elevenlabs/elevenlabs-python

    2,873在 GitHub 上查看↗

    This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of conversational AI agents. It enables the creation of lifelike spoken audio from text, the replication of human voices through custom cloning, and the deployment of real-time voice agents capable of interacting with external large language models. The library distinguishes itself through deep integration of conversational AI capabilities, including the design of agent personas and the execution of real-time actions via APIs. It supports professional-grade audio production thro

    Pythonartificial-intelligenceconversational-aitext-to-speech
    在 GitHub 上查看↗2,873
  • livekit/livekitlivekit 的头像

    livekit/livekit

    19,358在 GitHub 上查看↗

    LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with users through voice, video, and text. It provides a centralized, event-driven architecture to manage the entire lifecycle of automated participants, from initialization and session state management to graceful shutdown. By utilizing a selective forwarding unit, the platform efficiently routes media streams between participants and agents, ensuring low-latency communication and secure, token-based authentication for all connections. The platform distinguishes itself through it

    Gogolangmedia-serversfu
    在 GitHub 上查看↗19,358
查看 Speech To Speech 的所有 30 个替代方案→

常见问题解答

huggingface/speech-to-speech 是做什么的?

该项目是一个用于构建本地语音助手和实时音频流服务器的框架。它作为一个容器化推理引擎和多语言语音流水线,编排语音转文字 (STT)、语言模型和文字转语音 (TTS) 组件,将语音输入转换为语音输出。

huggingface/speech-to-speech 的主要功能有哪些?

huggingface/speech-to-speech 的主要功能包括:Voice Agents, Voice Assistants, Turn Event Emitters, Voice Activity Detection, Voice Pipelines, Voice Interaction Management, Model Provider Integrations, Real-Time Transcription。

huggingface/speech-to-speech 有哪些开源替代品?

huggingface/speech-to-speech 的开源替代品包括: getstream/vision-agents. vocodedev/vocode-core — Vocode-core is a framework for building real-time conversational AI voice agents. It serves as a conversational… elevenlabs/elevenlabs-python — This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of… livekit/livekit — LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with… livekit/agents — This project is a framework for developing multimodal AI agents that function as programmable participants in… pipecat-ai/pipecat — Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech…