awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
pluja avatar

pluja/whishper

0
View on GitHub↗
2,920 stars·167 forks·Svelte·agpl-3.0·13 viewswhishper-docs.pages.dev↗

Whishper

Whishper is a graphical user interface for transcribing audio and video files into text using the Whisper model. It serves as a speech-to-text tool and subtitle file generator that converts spoken content into editable text and timed subtitle formats.

The project features an integrated transcription and translation interface, allowing users to refine automated results and convert transcribed text into different languages. It includes a visual editor for correcting speech recognition errors, adjusting segment timecodes, and performing bilingual translation reviews.

The system handles the full transcription workflow, from retrieving media via remote URLs to exporting final data. Supported export formats include SRT, VTT, JSON, and plain text.

Features

  • Graphical User Interfaces - Provides a comprehensive graphical user interface for transcribing audio and video files using the Whisper model.
  • Speech to Text Transcription - Converts spoken content from uploaded or remote media files into text transcripts and subtitles.
  • Audio and Video File Transcription - Extracts speech from local media files to produce offline subtitles, plain text, and timestamp data.
  • Transcription Exporters - Saves transcriptions and subtitles into various formats such as plain text, JSON, VTT, and SRT.
  • Transcript Editors - Provides a visual editor for refining transcription text with segment splitting and timing adjustments.
  • Subtitle Segment Management - Enables precise control over the flow of subtitles by modifying individual text segment timecodes.
  • Language Translation Services - Translates transcribed text into multiple languages to improve accessibility of audio and video content.
  • Speech Recognition Engines - Uses an optimized local inference engine to convert audio to text while maintaining data privacy.
  • Transcript Refinement - Provides tools to correct speech recognition errors and refine segment timing for accurate transcripts.
  • Multi-Format Exports - Exports internal transcription data into multiple standardized formats including SRT, VTT, and JSON.
  • Timestamped Subtitle Generators - Generates timestamped subtitle files in SRT and VTT formats for video playback.
  • Media Text Digitization - Turns spoken content from uploaded files or remote URLs into editable and searchable text documents.
  • Time-Coded Segment Mapping - Organizes transcribed text into time-coded segments to synchronize subtitles with the audio track.
  • AI Translation Tools - Implements a workflow to convert transcribed spoken content into different languages using integrated translation engines.
  • Translation Editors - Offers an interface for reviewing and correcting automated translations against the original transcription.
  • Remote Media Fetching - Fetches audio and video files from remote URLs into a local buffer for processing.
  • Visual State Reconciliation - Implements real-time synchronization between the visual transcription editor and the underlying data state.
  • Bilingual Display Components - Provides a UI component that displays original and translated text segments side-by-side for linguistic verification.

Star history

Star history chart for pluja/whishperStar history chart for pluja/whishper

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does pluja/whishper do?

Whishper is a graphical user interface for transcribing audio and video files into text using the Whisper model. It serves as a speech-to-text tool and subtitle file generator that converts spoken content into editable text and timed subtitle formats.

What are the main features of pluja/whishper?

The main features of pluja/whishper are: Graphical User Interfaces, Speech to Text Transcription, Audio and Video File Transcription, Transcription Exporters, Transcript Editors, Subtitle Segment Management, Language Translation Services, Speech Recognition Engines.

Which projects share features with pluja/whishper?

Projects with overlapping indexed features include: cmusphinx/pocketsphinx — PocketSphinx is an offline speech recognition engine that converts raw audio from files or live microphone streams… kedreamix/linly-dubbing — Linly-Dubbing is an automated video dubbing pipeline designed for multilingual video localization. It converts spoken… jianchang512/stt — This project is a hardware-accelerated transcription server and offline subtitle generator. It functions as a… elevenlabs/elevenlabs-python — This Python SDK provides a comprehensive toolkit for synthetic audio generation, voice cloning, and the development of… agermanidis/autosub — Autosub is a command-line media processor and automatic subtitle generator that converts audio streams from video and… linyqh/narratoai — NarratoAI is an automated video production pipeline that uses large language models to generate scripts, voiceovers,…

Projects sharing features with Whishper

These projects share indexed features with Whishper. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • cmusphinx/pocketsphinxcmusphinx avatar

    cmusphinx/pocketsphinx

    4,276View on GitHub↗

    PocketSphinx is an offline speech recognition engine that converts raw audio from files or live microphone streams into written text without requiring a network connection. It functions as a speech-to-text library, a real-time transcription engine, and a voice command processor, capable of detecting and transcribing spoken commands from continuous audio streams with configurable acoustic and language models. The engine uses weighted finite-state transducers to represent acoustic, phonetic, and language models as a single search graph for efficient decoding. It employs fixed-point acoustic mod

    Ccpythonspeech-recognition
    View on GitHub↗4,276
  • kedreamix/linly-dubbingKedreamix avatar

    Kedreamix/Linly-Dubbing

    3,048View on GitHub↗

    Linly-Dubbing is an automated video dubbing pipeline designed for multilingual video localization. It converts spoken content in videos into another language by coordinating speech-to-text transcription, text translation, and text-to-speech synthesis. The system distinguishes itself through AI-driven lip synchronization and animation, which aligns facial expressions and mouth movements to the synthesized voiceover. It also utilizes audio source separation to isolate vocals from background music and noise, allowing for clean voice replacement while preserving original background audio. The br

    Jupyter Notebook
    View on GitHub↗3,048
  • jianchang512/sttjianchang512 avatar

    jianchang512/stt

    4,629View on GitHub↗

    This project is a hardware-accelerated transcription server and offline subtitle generator. It functions as a speech-to-text tool that converts audio and video files into plain text, JSON, and SRT subtitle formats using the Whisper model. The system operates as an OpenAI Audio API emulator, providing a local server that mimics a specific audio interface. This allows it to serve transcriptions to existing client configurations without requiring changes to the client software. The service utilizes GPU acceleration to increase voice recognition speed and includes utilities for hardware detectio

    Pythonspeechspeech-recognitionspeech-to-text
    View on GitHub↗4,629
  • agermanidis/autosubagermanidis avatar

    agermanidis/autosub

    4,197View on GitHub↗

    Autosub is a command-line media processor and automatic subtitle generator that converts audio streams from video and audio files into timed text overlays. It functions as an AI speech-to-text converter that uses OpenAI Whisper to generate synchronized subtitles. The tool includes a language translation pipeline to convert transcribed speech into target languages, enabling multilingual video captioning. It manages the process from audio-stream extraction to the serialization of final subtitle files for local storage. The system covers audio-to-text transcription, time-stamped text mapping, a

    Python
    View on GitHub↗4,197
Compare all 30 related projects→