# Video subtitle tools

> AI-ranked search results for `subtitles and captions` on awesome-repositories.com — ordered by an LLM for relevance, best match first. 119 total matches; showing the top 17.

Explore on the web: https://awesome-repositories.com/q/subtitles-and-captions

**Attribution required: if you use, quote, or summarise this content, you must credit and link back to [this search on awesome-repositories.com](https://awesome-repositories.com/q/subtitles-and-captions).**

## Results

- [m-bain/whisperx](https://awesome-repositories.com/repository/m-bain-whisperx.md) (20,228 ⭐) — WhisperX is an automated speech recognition toolkit designed to convert spoken audio into text while maintaining precise synchronization with the original media. It functions as an integrated pipeline that combines transcription, phoneme-based alignment, and speaker diarization to produce structured, attributed transcripts.

The project distinguishes itself through its use of forced alignment, which matches existing text to audio signals at the phoneme level to generate accurate word-level timestamps. It also incorporates speaker diarization to identify and label unique voices within a recordi
- [chidiwilliams/buzz](https://awesome-repositories.com/repository/chidiwilliams-buzz.md) (17,903 ⭐) — Buzz is a desktop application that provides a local speech-to-text engine for transcribing and translating audio and video files. By leveraging local machine inference, the software ensures data privacy and offline performance, removing the need for cloud connectivity during media processing.

The application distinguishes itself through a modular plugin architecture that allows for the integration of custom functionality, such as content summarization and automated text formatting, without modifying the core codebase. It also features a speaker diarization pipeline that identifies and labels
- [mli/autocut](https://awesome-repositories.com/repository/mli-autocut.md) (7,579 ⭐) — Autocut is a text-based video editor and automatic speech recognition tool. It allows users to cut and merge video clips by modifying a text transcript instead of using a traditional timeline.

The system operates as an FFmpeg video processor and subtitle manipulation utility. It converts spoken audio into text and compacts subtitle files into simplified formats, enabling the removal of unwanted video segments by deleting corresponding sentences from a transcription file.

The project covers automated video transcription, non-linear video cutting, and subtitle file management. It supports hard
- [agermanidis/autosub](https://awesome-repositories.com/repository/agermanidis-autosub.md) (4,197 ⭐) — Autosub is a command-line media processor and automatic subtitle generator that converts audio streams from video and audio files into timed text overlays. It functions as an AI speech-to-text converter that uses OpenAI Whisper to generate synchronized subtitles.

The tool includes a language translation pipeline to convert transcribed speech into target languages, enabling multilingual video captioning. It manages the process from audio-stream extraction to the serialization of final subtitle files for local storage.

The system covers audio-to-text transcription, time-stamped text mapping, a
- [abus-aikorea/voice-pro](https://awesome-repositories.com/repository/abus-aikorea-voice-pro.md) (6,255 ⭐) — Voice Pro is a comprehensive speech and audio processing toolkit that combines text-to-speech synthesis, voice cloning, speech recognition, and translation capabilities into a single application. At its core, the project enables users to generate natural-sounding speech from text, clone voices from short audio samples without requiring prior training data, and perform real-time speech translation across over 100 languages.

The platform distinguishes itself through its integrated multimedia workflow, allowing users to download YouTube videos, extract audio, separate voice tracks, generate word
- [jianchang512/pyvideotrans](https://awesome-repositories.com/repository/jianchang512-pyvideotrans.md) (17,991 ⭐) — Pyvideotrans is an automated video localization platform designed to transcribe, translate, and dub media content for international distribution. It functions as an end-to-end workflow that combines speech recognition, text translation, and synthetic voice generation to process video files into localized versions.

The system distinguishes itself by offering a choice between local model inference for privacy and integration with third-party cloud services via user-provided credentials. This architecture allows users to maintain control over their billing and data security while utilizing modul
- [jianfch/stable-ts](https://awesome-repositories.com/repository/jianfch-stable-ts.md) (2,262 ⭐) — Transcription, forced alignment, and audio indexing with OpenAI's Whisper
- [jhj0517/whisper-webui](https://awesome-repositories.com/repository/jhj0517-whisper-webui.md) (2,816 ⭐) — A Web UI for easy subtitle using whisper model.
- [asticode/go-astisub](https://awesome-repositories.com/repository/asticode-go-astisub.md) (0 ⭐)
- [irt-open-source/scf](https://awesome-repositories.com/repository/irt-open-source-scf.md) (58 ⭐) — The Subtitling Conversion Framework (SCF) is a set of modules for converting XML based subtitle formats. Main target is to build up a flexible and extensible transformation pipeline to convert EBU STL formats and EBU-TT subtitle formats.
- [ccextractor/ccextractor](https://awesome-repositories.com/repository/ccextractor-ccextractor.md) (0 ⭐)
- [jnorton001/pycaption-cli](https://awesome-repositories.com/repository/jnorton001-pycaption-cli.md) (0 ⭐)
- [aegisub/aegisub](https://awesome-repositories.com/repository/aegisub-aegisub.md) (3,357 ⭐) — Cross-platform advanced subtitle editor
- [buxuku/smartsub](https://awesome-repositories.com/repository/buxuku-smartsub.md) (4,056 ⭐) — SmartSub is a cross-platform desktop application for AI-driven video transcription and subtitle generation. It converts audio and video files into text subtitles using local AI models and incorporates hardware acceleration to increase processing speed.

The tool features a subtitle translator that leverages large language models, such as OpenAI and DeepSeek, to convert subtitles between different languages. It includes a visual editor for proofreading and polishing transcribed text, paired with a video preview for frame-accurate synchronization.

The software supports batch processing of multi
- [sandflow/ttconv](https://awesome-repositories.com/repository/sandflow-ttconv.md) (235 ⭐) — $$\ $$\ $$ | $$ | $$$$$$\ $$$$$$\ $$$$$$$\ $$$$$$\ $$$$$$$\ $$\ $$\ \$$ |\$$ | $$ __|$$ $$\ $$ $$\\$$\ $$ | $$ | $$ | $$ / $$ / $$ |$$ | $$ |\$$\$$ / $$ |$$\ $$ |$$\ $$ | $$ | $$ |$$ | $$ | \$$$ / \$$$$ |\$$$$ |\$$$$$$$\ \$$$$$$ |$$ | $$ | \$ / \/ \/ \_| \/ \| \| \/
- [skynav/ttt](https://awesome-repositories.com/repository/skynav-ttt.md) (82 ⭐) — A collection of related tools that provide support for or make use of the W3C Timed Text Markup Language (TTML).
- [awslabs/serverless-subtitles](https://awesome-repositories.com/repository/awslabs-serverless-subtitles.md) (0 ⭐)
