For subtitles and captions, the first results are m-bain/whisperx (WhisperX provides automatic speech recognition and precise timing synchronization for generating captions, serving as a powerful pipeline tool even though it lacks a full subtitle editing interface or built-in video player), chidiwilliams/buzz (Buzz is a desktop application that provides local speech-to-text transcription and translation for audio and video files, fitting the category through its focus on generating subtitles despite lacking a dedicated editing interface) and mli/autocut (Autocut provides automatic speech recognition and transcript-driven subtitle processing to enable text-based video editing and manipulation, fulfilling several key aspects of the search). agermanidis/autosub and abus-aikorea/voice-pro round out the shortlist. Compare the match explanations and check the project documentation against your requirements.
Hand-picked open-source video subtitle and caption tools, ranked by GitHub stars and activity. Compare the top alternatives and find the best fit.
WhisperX is an automated speech recognition toolkit designed to convert spoken audio into text while maintaining precise synchronization with the original media. It functions as an integrated pipeline that combines transcription, phoneme-based alignment, and speaker diarization to produce structured, attributed transcripts. The project distinguishes itself through its use of forced alignment, which matches existing text to audio signals at the phoneme level to generate accurate word-level timestamps. It also incorporates speaker diarization to identify and label unique voices within a recordi
WhisperX provides automatic speech recognition and precise timing synchronization for generating captions, serving as a powerful pipeline tool even though it lacks a full subtitle editing interface or built-in video player.
Buzz is a desktop application that provides a local speech-to-text engine for transcribing and translating audio and video files. By leveraging local machine inference, the software ensures data privacy and offline performance, removing the need for cloud connectivity during media processing. The application distinguishes itself through a modular plugin architecture that allows for the integration of custom functionality, such as content summarization and automated text formatting, without modifying the core codebase. It also features a speaker diarization pipeline that identifies and labels
Buzz is a desktop application that provides local speech-to-text transcription and translation for audio and video files, fitting the category through its focus on generating subtitles despite lacking a dedicated editing interface.
Autocut is a text-based video editor and automatic speech recognition tool. It allows users to cut and merge video clips by modifying a text transcript instead of using a traditional timeline. The system operates as an FFmpeg video processor and subtitle manipulation utility. It converts spoken audio into text and compacts subtitle files into simplified formats, enabling the removal of unwanted video segments by deleting corresponding sentences from a transcription file. The project covers automated video transcription, non-linear video cutting, and subtitle file management. It supports hard
Autocut provides automatic speech recognition and transcript-driven subtitle processing to enable text-based video editing and manipulation, fulfilling several key aspects of the search.
Autosub is a command-line media processor and automatic subtitle generator that converts audio streams from video and audio files into timed text overlays. It functions as an AI speech-to-text converter that uses OpenAI Whisper to generate synchronized subtitles. The tool includes a language translation pipeline to convert transcribed speech into target languages, enabling multilingual video captioning. It manages the process from audio-stream extraction to the serialization of final subtitle files for local storage. The system covers audio-to-text transcription, time-stamped text mapping, a
Autosub is a command-line automatic subtitle generator that uses speech recognition to transcribe and translate audio into timed subtitle files, though it lacks a visual editing interface.
Voice Pro is a comprehensive speech and audio processing toolkit that combines text-to-speech synthesis, voice cloning, speech recognition, and translation capabilities into a single application. At its core, the project enables users to generate natural-sounding speech from text, clone voices from short audio samples without requiring prior training data, and perform real-time speech translation across over 100 languages. The platform distinguishes itself through its integrated multimedia workflow, allowing users to download YouTube videos, extract audio, separate voice tracks, generate word
Voice Pro is an audio and speech processing application that includes automated video subtitling, speech-to-text engines, and timestamped subtitle generation, fitting the search for video subtitle tools despite its broader focus on voice cloning and TTS.
Pyvideotrans is an automated video localization platform designed to transcribe, translate, and dub media content for international distribution. It functions as an end-to-end workflow that combines speech recognition, text translation, and synthetic voice generation to process video files into localized versions. The system distinguishes itself by offering a choice between local model inference for privacy and integration with third-party cloud services via user-provided credentials. This architecture allows users to maintain control over their billing and data security while utilizing modul
Pyvideotrans is an automated video localization platform that handles speech-to-text transcription and subtitle generation, though it leans heavily toward translation and dubbing rather than a dedicated subtitle-editing interface.
Transcription, forced alignment, and audio indexing with OpenAI's Whisper
This library uses Whisper for forced alignment and transcription to generate precise subtitle timings, fitting the core need for automated subtitle generation despite lacking an editing interface.
A Web UI for easy subtitle using whisper model.
This repository provides a web interface for generating subtitles using the Whisper speech recognition model, fitting the category through its focus on subtitle creation.
This Go library provides programmatic tools for reading, writing, converting, and syncing subtitles across various formats, making it a solid backend building block for custom caption workflows.
The Subtitling Conversion Framework (SCF) is a set of modules for converting XML based subtitle formats. Main target is to build up a flexible and extensible transformation pipeline to convert EBU STL formats and EBU-TT subtitle formats.
The Subtitling Conversion Framework is a specialized module set for converting XML-based subtitle formats like EBU STL and EBU-TT, which fulfills the subtitle format conversion requirement even though it lacks an editing interface or video player.
CCExtractor is a command-line tool dedicated to extracting and processing closed captions from video files, though it lacks a full interactive subtitle editing interface or automatic speech recognition.
This command-line utility handles subtitle format conversion between various caption standards, fitting the targeted functionality even though it lacks an editing interface or speech recognition.
Cross-platform advanced subtitle editor
Aegisub is a cross-platform subtitle editor featuring precise timing synchronization and an editing interface, making it a fitting tool for this category despite lacking automatic speech recognition and conversion features out of the box.
SmartSub is a cross-platform desktop application for AI-driven video transcription and subtitle generation. It converts audio and video files into text subtitles using local AI models and incorporates hardware acceleration to increase processing speed. The tool features a subtitle translator that leverages large language models, such as OpenAI and DeepSeek, to convert subtitles between different languages. It includes a visual editor for proofreading and polishing transcribed text, paired with a video preview for frame-accurate synchronization. The software supports batch processing of multi
SmartSub is a desktop application providing AI-driven video transcription, translation, and a visual editor with video preview for timeline synchronization, matching most of the required subtitle and closed captioning tools.
$$\ $$\ $$ | $$ | $$$$$$\ $$$$$$\ $$$$$$$\ $$$$$$\ $$$$$$$\ $$\ $$\ \$$ |\$$ | $$ |$$ $$\ $$ $$\\$$\ $$ | $$ | $$ | $$ / $$ / $$ |$$ | $$ |\$$\$$ / $$ |$$\ $$ |$$\ $$ | $$ | $$ |$$ | $$ | \$$$ / \$$$$ |\$$$$ |\$$$$$$$\ \$$$$$$ |$$ | $$ | \$ / \/ \/ \_| \/ \| \| \/
This Python library handles subtitle format conversion and processing, making it a fitting tool for subtitle manipulation despite lacking an editing interface or speech recognition.
A collection of related tools that provide support for or make use of the W3C Timed Text Markup Language (TTML).
This Java library provides tools for working with W3C Timed Text Markup Language (TTML) subtitles, making it a relevant option for format conversion and manipulation, although it lacks an interactive editing interface or built-in speech recognition.
This serverless project handles subtitle generation and management pipelines in the cloud, fitting the requested category despite lacking a visible description or detailed feature list.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| m-bain/whisperx | 20.2K | Python | bsd-2-clause | |
| chidiwilliams/buzz | 17.9K | Python | mit | |
| mli/autocut |
| 7.6K |
| Python |
| apache-2.0 |
| agermanidis/autosub | 4.2K | Python | MIT |
| abus-aikorea/voice-pro | 6.3K | Python | gpl-3.0 |
| jianchang512/pyvideotrans | 18K | Python | GPL-3.0 |
| jianfch/stable-ts | 2.3K | Python | MIT |
| jhj0517/whisper-webui | 2.8K | Python | Apache-2.0 |
| asticode/go-astisub | 0 | — | — | — |
| irt-open-source/scf | 58 | XSLT | Apache-2.0 |