awesome-repositories.com
博客
MCP
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目MCP 服务器关于排名机制媒体报道
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
chidiwilliams avatar

chidiwilliams/buzz

0
View on GitHub↗
17,903 星标·1,313 分支·Python·mit·31 次浏览chidiwilliams.github.io/buzz↗

Buzz

Buzz is a desktop application that provides a local speech-to-text engine for transcribing and translating audio and video files. By leveraging local machine inference, the software ensures data privacy and offline performance, removing the need for cloud connectivity during media processing.

The application distinguishes itself through a modular plugin architecture that allows for the integration of custom functionality, such as content summarization and automated text formatting, without modifying the core codebase. It also features a speaker diarization pipeline that identifies and labels individual voices within recordings to improve the readability and organization of generated transcripts.

The system supports automated media processing by monitoring specific directories for new files, enabling users to trigger transcription or translation workflows as soon as assets are detected. Users can export results into various standard formats, including plain text and subtitle files, while utilizing hardware acceleration to increase processing speeds for large media files.

Features

  • Speech-to-Text Engines - Provides a local speech-to-text engine that leverages hardware acceleration and speaker diarization.
  • Audio Transcription - Converts audio and video files into written text locally to ensure complete data privacy.
  • Speech-to-Text Utilities - Provides a local speech-to-text engine for transcribing audio and video files offline.
  • Transcription Tools - Converts audio and video files into text using local machine processing to ensure privacy and offline performance.
  • Local AI Inference - Executes machine learning models locally to ensure data privacy and offline performance.
  • Multilingual Speech Translation - Converts spoken language from media into different languages using local machine inference.
  • Speech-to-Text Translation - Translates spoken language from media files into different languages using local processing.
  • Speaker Diarization - Analyzes audio to distinguish and label individual speakers within transcripts.
  • Desktop Applications - Cross-platform tool for local audio transcription and translation.
  • Voice Dictation - Offline audio transcription and translation using AI models.
  • Voice To Text - Listed in the “Voice To Text” section of the Awesome Mac awesome list.
  • Speech Recognition - Desktop application for speech recognition and subtitle generation.
  • Media Automation - Automates transcription and translation tasks by monitoring directories for new media assets.
  • Modular Plugin Architectures - Provides a modular plugin architecture that allows for the integration of custom functionality like summarization and formatting without modifying the core codebase.
  • Plugin Architectures - Enables extending core functionality through modular plugins for tasks like summarization.
  • Subtitle Management Systems - Exports transcripts into standard subtitle and web video track formats.
  • Hardware Acceleration - Offloads intensive transcription tasks to local graphics hardware to increase processing speed.

Star 历史

chidiwilliams/buzz 的 Star 历史图表chidiwilliams/buzz 的 Star 历史图表

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Start searching with AI

常见问题解答

chidiwilliams/buzz 是做什么的?

Buzz is a desktop application that provides a local speech-to-text engine for transcribing and translating audio and video files. By leveraging local machine inference, the software ensures data privacy and offline performance, removing the need for cloud connectivity during media processing.

chidiwilliams/buzz 的主要功能有哪些?

chidiwilliams/buzz 的主要功能包括:Speech-to-Text Engines, Audio Transcription, Speech-to-Text Utilities, Transcription Tools, Local AI Inference, Multilingual Speech Translation, Speech-to-Text Translation, Speaker Diarization。

chidiwilliams/buzz 有哪些开源替代品?

chidiwilliams/buzz 的开源替代品包括: pipecat-ai/pipecat — Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech… cjpais/handy — Handy is a local speech-to-text automation tool designed to convert spoken audio into text and inject it directly into… livekit/livekit — LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with… k2-fsa/sherpa-onnx — Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device… argmaxinc/whisperkit. jamiepine/voicebox — Voicebox is a local speech processing system that provides text-to-speech generation, speech-to-text transcription,…

Buzz 的开源替代方案

相似的开源项目,按与 Buzz 的功能重合度排序。
  • pipecat-ai/pipecatpipecat-ai 的头像

    pipecat-ai/pipecat

    12,846在 GitHub 上查看↗

    Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech systems. It utilizes a frame-based data pipeline to route audio, video, and text through a modular sequence of processors, enabling the orchestration of low-latency conversational AI. The project is distinguished by its ability to coordinate complex multimodal services, including speech-to-text, language models, and text-to-speech, within a single pipeline. It features semantic voice activity detection for natural turn-taking, state-machine conversation flows for dialogue manag

    Pythonaichatbot-frameworkchatbots
    在 GitHub 上查看↗12,846
  • cjpais/handycjpais 的头像

    cjpais/Handy

    15,515在 GitHub 上查看↗

    Handy is a local speech-to-text automation tool designed to convert spoken audio into text and inject it directly into active desktop applications. By running machine learning models entirely on the host hardware, it provides a private, offline-first environment for dictation and command execution. The system functions as a background service that manages microphone input, transcription state, and text output, enabling hands-free typing across various software environments. The project distinguishes itself through a modular pipeline that integrates local language models for post-transcription

    Rustaccessibilitycross-platformspeech-to-text
    在 GitHub 上查看↗15,515
  • livekit/livekitlivekit 的头像

    livekit/livekit

    19,358在 GitHub 上查看↗

    LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with users through voice, video, and text. It provides a centralized, event-driven architecture to manage the entire lifecycle of automated participants, from initialization and session state management to graceful shutdown. By utilizing a selective forwarding unit, the platform efficiently routes media streams between participants and agents, ensuring low-latency communication and secure, token-based authentication for all connections. The platform distinguishes itself through it

    Gogolangmedia-serversfu
    在 GitHub 上查看↗19,358
  • k2-fsa/sherpa-onnxk2-fsa 的头像

    k2-fsa/sherpa-onnx

    13,017在 GitHub 上查看↗

    Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device voice synthesis tool, and a speaker identification framework. It is designed as a cross-platform speech API that enables speech-to-text, text-to-speech, and speaker verification tasks to be executed locally on a device without requiring network access. The project is distinguished by its ability to perform zero-shot voice cloning and speaker diarization on-device. It supports a wide range of hardware accelerations, including GPU and various NPU architectures, and provides a Web

    C++aarch64androidarm32
    在 GitHub 上查看↗13,017
  • 查看 Buzz 的所有 30 个替代方案→