awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
chidiwilliams avatar

chidiwilliams/buzz

0
View on GitHub↗
17,903 stars·1,313 forks·Python·mit·47 viewschidiwilliams.github.io/buzz↗

Buzz

Buzz is a desktop application that provides a local speech-to-text engine for transcribing and translating audio and video files. By leveraging local machine inference, the software ensures data privacy and offline performance, removing the need for cloud connectivity during media processing.

The application distinguishes itself through a modular plugin architecture that allows for the integration of custom functionality, such as content summarization and automated text formatting, without modifying the core codebase. It also features a speaker diarization pipeline that identifies and labels individual voices within recordings to improve the readability and organization of generated transcripts.

The system supports automated media processing by monitoring specific directories for new files, enabling users to trigger transcription or translation workflows as soon as assets are detected. Users can export results into various standard formats, including plain text and subtitle files, while utilizing hardware acceleration to increase processing speeds for large media files.

Features

  • Speech-to-Text Engines - Provides a local speech-to-text engine that leverages hardware acceleration and speaker diarization.
  • Audio Transcription - Converts audio and video files into written text locally to ensure complete data privacy.
  • Speech-to-Text Utilities - Provides a local speech-to-text engine for transcribing audio and video files offline.
  • Transcription Tools - Converts audio and video files into text using local machine processing to ensure privacy and offline performance.
  • Local AI Inference - Executes machine learning models locally to ensure data privacy and offline performance.
  • Multilingual Speech Translation - Converts spoken language from media into different languages using local machine inference.
  • Speech-to-Text Translation - Translates spoken language from media files into different languages using local processing.
  • Speaker Diarization - Analyzes audio to distinguish and label individual speakers within transcripts.
  • Desktop Applications - Cross-platform tool for local audio transcription and translation.
  • Voice Dictation - Offline audio transcription and translation using AI models.
  • Voice To Text - Listed in the “Voice To Text” section of the Awesome Mac awesome list.
  • Speech Recognition - Desktop application for speech recognition and subtitle generation.
  • Media Automation - Automates transcription and translation tasks by monitoring directories for new media assets.
  • Modular Plugin Architectures - Provides a modular plugin architecture that allows for the integration of custom functionality like summarization and formatting without modifying the core codebase.
  • Plugin Architectures - Enables extending core functionality through modular plugins for tasks like summarization.
  • Subtitle Management Systems - Exports transcripts into standard subtitle and web video track formats.
  • Hardware Acceleration - Offloads intensive transcription tasks to local graphics hardware to increase processing speed.

Star history

Star history chart for chidiwilliams/buzzStar history chart for chidiwilliams/buzz

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does chidiwilliams/buzz do?

Buzz is a desktop application that provides a local speech-to-text engine for transcribing and translating audio and video files. By leveraging local machine inference, the software ensures data privacy and offline performance, removing the need for cloud connectivity during media processing.

What are the main features of chidiwilliams/buzz?

The main features of chidiwilliams/buzz are: Speech-to-Text Engines, Audio Transcription, Speech-to-Text Utilities, Transcription Tools, Local AI Inference, Multilingual Speech Translation, Speech-to-Text Translation, Speaker Diarization.

Which projects share features with chidiwilliams/buzz?

Projects with overlapping indexed features include: pipecat-ai/pipecat — Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech… cjpais/handy — Handy is a local speech-to-text automation tool designed to convert spoken audio into text and inject it directly into… livekit/livekit — LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with… k2-fsa/sherpa-onnx — Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device… argmaxinc/whisperkit. jamiepine/voicebox — Voicebox is a local speech processing system that provides text-to-speech generation, speech-to-text transcription,…

Projects sharing features with Buzz

These projects share indexed features with Buzz. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • pipecat-ai/pipecatpipecat-ai avatar

    pipecat-ai/pipecat

    12,846View on GitHub↗

    Pipecat is a framework and software development kit for building real-time multimodal AI agents and speech-to-speech systems. It utilizes a frame-based data pipeline to route audio, video, and text through a modular sequence of processors, enabling the orchestration of low-latency conversational AI. The project is distinguished by its ability to coordinate complex multimodal services, including speech-to-text, language models, and text-to-speech, within a single pipeline. It features semantic voice activity detection for natural turn-taking, state-machine conversation flows for dialogue manag

    Pythonaichatbot-frameworkchatbots
    View on GitHub↗12,846
  • cjpais/handycjpais avatar

    cjpais/Handy

    15,515View on GitHub↗

    Handy is a local speech-to-text automation tool designed to convert spoken audio into text and inject it directly into active desktop applications. By running machine learning models entirely on the host hardware, it provides a private, offline-first environment for dictation and command execution. The system functions as a background service that manages microphone input, transcription state, and text output, enabling hands-free typing across various software environments. The project distinguishes itself through a modular pipeline that integrates local language models for post-transcription

    Rustaccessibilitycross-platformspeech-to-text
    View on GitHub↗15,515
  • livekit/livekitlivekit avatar

    livekit/livekit

    19,358View on GitHub↗

    LiveKit is a comprehensive framework for building and orchestrating real-time, multimodal AI agents that interact with users through voice, video, and text. It provides a centralized, event-driven architecture to manage the entire lifecycle of automated participants, from initialization and session state management to graceful shutdown. By utilizing a selective forwarding unit, the platform efficiently routes media streams between participants and agents, ensuring low-latency communication and secure, token-based authentication for all connections. The platform distinguishes itself through it

    Gogolangmedia-serversfu
    View on GitHub↗19,358
  • k2-fsa/sherpa-onnxk2-fsa avatar

    k2-fsa/sherpa-onnx

    13,017View on GitHub↗

    Sherpa-ONNX is an ONNX-based speech processing toolkit that provides a local speech recognition engine, an on-device voice synthesis tool, and a speaker identification framework. It is designed as a cross-platform speech API that enables speech-to-text, text-to-speech, and speaker verification tasks to be executed locally on a device without requiring network access. The project is distinguished by its ability to perform zero-shot voice cloning and speaker diarization on-device. It supports a wide range of hardware accelerations, including GPU and various NPU architectures, and provides a Web

    C++aarch64androidarm32
    View on GitHub↗13,017
Compare all 30 related projects→