awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com

Video subtitle tools

Ranking updated Sep 7, 2026

For subtitles and captions, the first results are m-bain/whisperx (WhisperX provides automatic speech recognition and precise timing synchronization for generating captions, serving as a powerful pipeline tool even though it lacks a full subtitle editing interface or built-in video player), chidiwilliams/buzz (Buzz is a desktop application that provides local speech-to-text transcription and translation for audio and video files, fitting the category through its focus on generating subtitles despite lacking a dedicated editing interface) and mli/autocut (Autocut provides automatic speech recognition and transcript-driven subtitle processing to enable text-based video editing and manipulation, fulfilling several key aspects of the search). agermanidis/autosub and abus-aikorea/voice-pro round out the shortlist. Compare the match explanations and check the project documentation against your requirements.

Hand-picked open-source video subtitle and caption tools, ranked by GitHub stars and activity. Compare the top alternatives and find the best fit.

Video subtitle tools

Find the best repos with AI.We'll search the best matching repositories with AI.
  • m-bain/whisperxm-bain avatar

    m-bain/whisperX

    20,228View on GitHub↗

    WhisperX is an automated speech recognition toolkit designed to convert spoken audio into text while maintaining precise synchronization with the original media. It functions as an integrated pipeline that combines transcription, phoneme-based alignment, and speaker diarization to produce structured, attributed transcripts. The project distinguishes itself through its use of forced alignment, which matches existing text to audio signals at the phoneme level to generate accurate word-level timestamps. It also incorporates speaker diarization to identify and label unique voices within a recordi

    WhisperX provides automatic speech recognition and precise timing synchronization for generating captions, serving as a powerful pipeline tool even though it lacks a full subtitle editing interface or built-in video player.

    PythonAutomatic Speech RecognitionForced AlignmentAudio Transcription
    View on GitHub↗20,228
  • chidiwilliams/buzzchidiwilliams avatar

    chidiwilliams/buzz

    17,903View on GitHub↗

    Buzz is a desktop application that provides a local speech-to-text engine for transcribing and translating audio and video files. By leveraging local machine inference, the software ensures data privacy and offline performance, removing the need for cloud connectivity during media processing. The application distinguishes itself through a modular plugin architecture that allows for the integration of custom functionality, such as content summarization and automated text formatting, without modifying the core codebase. It also features a speaker diarization pipeline that identifies and labels

    Buzz is a desktop application that provides local speech-to-text transcription and translation for audio and video files, fitting the category through its focus on generating subtitles despite lacking a dedicated editing interface.

    PythonAudio TranscriptionSpeech-to-Text EnginesSpeech-to-Text Utilities
    View on GitHub↗17,903
  • mli/autocutmli avatar

    mli/autocut

    7,579View on GitHub↗

    Autocut is a text-based video editor and automatic speech recognition tool. It allows users to cut and merge video clips by modifying a text transcript instead of using a traditional timeline. The system operates as an FFmpeg video processor and subtitle manipulation utility. It converts spoken audio into text and compacts subtitle files into simplified formats, enabling the removal of unwanted video segments by deleting corresponding sentences from a transcription file. The project covers automated video transcription, non-linear video cutting, and subtitle file management. It supports hard

    Autocut provides automatic speech recognition and transcript-driven subtitle processing to enable text-based video editing and manipulation, fulfilling several key aspects of the search.

    PythonAutomatic Speech RecognitionSubtitle Format Converters
    View on GitHub↗7,579
  • agermanidis/autosubagermanidis avatar

    agermanidis/autosub

    4,197View on GitHub↗

    Autosub is a command-line media processor and automatic subtitle generator that converts audio streams from video and audio files into timed text overlays. It functions as an AI speech-to-text converter that uses OpenAI Whisper to generate synchronized subtitles. The tool includes a language translation pipeline to convert transcribed speech into target languages, enabling multilingual video captioning. It manages the process from audio-stream extraction to the serialization of final subtitle files for local storage. The system covers audio-to-text transcription, time-stamped text mapping, a

    Autosub is a command-line automatic subtitle generator that uses speech recognition to transcribe and translate audio into timed subtitle files, though it lacks a visual editing interface.

    PythonSpeech Recognition APIsWhisper-Based Engines
    View on GitHub↗4,197
  • abus-aikorea/voice-proabus-aikorea avatar

    abus-aikorea/voice-pro

    6,255View on GitHub↗

    Voice Pro is a comprehensive speech and audio processing toolkit that combines text-to-speech synthesis, voice cloning, speech recognition, and translation capabilities into a single application. At its core, the project enables users to generate natural-sounding speech from text, clone voices from short audio samples without requiring prior training data, and perform real-time speech translation across over 100 languages. The platform distinguishes itself through its integrated multimedia workflow, allowing users to download YouTube videos, extract audio, separate voice tracks, generate word

    Voice Pro is an audio and speech processing application that includes automated video subtitling, speech-to-text engines, and timestamped subtitle generation, fitting the search for video subtitle tools despite its broader focus on voice cloning and TTS.

    PythonSpeech-to-Text Engines
    View on GitHub↗6,255
  • jianchang512/pyvideotransjianchang512 avatar

    jianchang512/pyvideotrans

    17,991View on GitHub↗

    Pyvideotrans is an automated video localization platform designed to transcribe, translate, and dub media content for international distribution. It functions as an end-to-end workflow that combines speech recognition, text translation, and synthetic voice generation to process video files into localized versions. The system distinguishes itself by offering a choice between local model inference for privacy and integration with third-party cloud services via user-provided credentials. This architecture allows users to maintain control over their billing and data security while utilizing modul

    Pyvideotrans is an automated video localization platform that handles speech-to-text transcription and subtitle generation, though it leans heavily toward translation and dubbing rather than a dedicated subtitle-editing interface.

    PythonSpeech Transcription
    View on GitHub↗17,991
  • jianfch/stable-tsjianfch avatar

    jianfch/stable-ts

    2,262View on GitHub↗

    Transcription, forced alignment, and audio indexing with OpenAI's Whisper

    This library uses Whisper for forced alignment and transcription to generate precise subtitle timings, fitting the core need for automated subtitle generation despite lacking an editing interface.

    PythonSubtitles and Captions
    View on GitHub↗2,262
  • jhj0517/whisper-webuijhj0517 avatar

    jhj0517/Whisper-WebUI

    2,816View on GitHub↗

    A Web UI for easy subtitle using whisper model.

    This repository provides a web interface for generating subtitles using the Whisper speech recognition model, fitting the category through its focus on subtitle creation.

    PythonSubtitles and Captions
    View on GitHub↗2,816
  • asticode/go-astisubA

    asticode/go-astisub

    0View on GitHub↗

    This Go library provides programmatic tools for reading, writing, converting, and syncing subtitles across various formats, making it a solid backend building block for custom caption workflows.

    Media ProcessingSubtitles and CaptionsVideo and Subtitle Handling
    View on GitHub↗0
  • irt-open-source/scfIrt-Open-Source avatar

    Irt-Open-Source/scf

    58View on GitHub↗

    The Subtitling Conversion Framework (SCF) is a set of modules for converting XML based subtitle formats. Main target is to build up a flexible and extensible transformation pipeline to convert EBU STL formats and EBU-TT subtitle formats.

    The Subtitling Conversion Framework is a specialized module set for converting XML-based subtitle formats like EBU STL and EBU-TT, which fulfills the subtitle format conversion requirement even though it lacks an editing interface or video player.

    XSLTSubtitles and CaptionsSubtitling
    View on GitHub↗58
  • ccextractor/ccextractorC

    CCExtractor/ccextractor

    0View on GitHub↗

    CCExtractor is a command-line tool dedicated to extracting and processing closed captions from video files, though it lacks a full interactive subtitle editing interface or automatic speech recognition.

    Subtitles and Captions
    View on GitHub↗0
  • jnorton001/pycaption-cliJ

    jnorton001/pycaption-cli

    0View on GitHub↗

    This command-line utility handles subtitle format conversion between various caption standards, fitting the targeted functionality even though it lacks an editing interface or speech recognition.

    Subtitles and Captions
    View on GitHub↗0
  • aegisub/aegisubAegisub avatar

    Aegisub/Aegisub

    3,357View on GitHub↗

    Cross-platform advanced subtitle editor

    Aegisub is a cross-platform subtitle editor featuring precise timing synchronization and an editing interface, making it a fitting tool for this category despite lacking automatic speech recognition and conversion features out of the box.

    C++Audio and VideoAudio Video Tools
    View on GitHub↗3,357
  • buxuku/smartsubbuxuku avatar

    buxuku/SmartSub

    4,056View on GitHub↗

    SmartSub is a cross-platform desktop application for AI-driven video transcription and subtitle generation. It converts audio and video files into text subtitles using local AI models and incorporates hardware acceleration to increase processing speed. The tool features a subtitle translator that leverages large language models, such as OpenAI and DeepSeek, to convert subtitles between different languages. It includes a visual editor for proofreading and polishing transcribed text, paired with a video preview for frame-accurate synchronization. The software supports batch processing of multi

    SmartSub is a desktop application providing AI-driven video transcription, translation, and a visual editor with video preview for timeline synchronization, matching most of the required subtitle and closed captioning tools.

    TypeScriptAutomated Subtitle GeneratorsAccelerated TranscriptionsAudio and Video File Transcription
    View on GitHub↗4,056
  • sandflow/ttconvsandflow avatar

    sandflow/ttconv

    235View on GitHub↗

    $$\ $$\ $$ | $$ | $$$$$$\ $$$$$$\ $$$$$$$\ $$$$$$\ $$$$$$$\ $$\ $$\ \$$ |\$$ | $$ |$$ $$\ $$ $$\\$$\ $$ | $$ | $$ | $$ / $$ / $$ |$$ | $$ |\$$\$$ / $$ |$$\ $$ |$$\ $$ | $$ | $$ |$$ | $$ | \$$$ / \$$$$ |\$$$$ |\$$$$$$$\ \$$$$$$ |$$ | $$ | \$ / \/ \/ \_| \/ \| \| \/

    This Python library handles subtitle format conversion and processing, making it a fitting tool for subtitle manipulation despite lacking an editing interface or speech recognition.

    PythonSubtitles and CaptionsSubtitling
    View on GitHub↗235
  • skynav/tttskynav avatar

    skynav/ttt

    82View on GitHub↗

    A collection of related tools that provide support for or make use of the W3C Timed Text Markup Language (TTML).

    This Java library provides tools for working with W3C Timed Text Markup Language (TTML) subtitles, making it a relevant option for format conversion and manipulation, although it lacks an interactive editing interface or built-in speech recognition.

    JavaSubtitling
    View on GitHub↗82
  • awslabs/serverless-subtitlesA

    awslabs/serverless-subtitles

    0View on GitHub↗

    This serverless project handles subtitle generation and management pipelines in the cloud, fitting the requested category despite lacking a visible description or detailed feature list.

    Subtitle and Caption Management
    View on GitHub↗0
Compare the top 10 at a glance
RepositoryStarsLanguageLicenseLast push
m-bain/whisperx20.2KPythonbsd-2-clauseFeb 19, 2026
chidiwilliams/buzz17.9KPythonmitFeb 19, 2026
mli/autocut
7.6K
Python
apache-2.0
Oct 5, 2024
agermanidis/autosub4.2KPythonMITMar 22, 2024
abus-aikorea/voice-pro6.3KPythongpl-3.0Dec 5, 2025
jianchang512/pyvideotrans18KPythonGPL-3.0Jun 16, 2026
jianfch/stable-ts2.3KPythonMITMay 30, 2026
jhj0517/whisper-webui2.8KPythonApache-2.0Dec 29, 2025
asticode/go-astisub0———
irt-open-source/scf58XSLTApache-2.0Nov 16, 2020

Related searches

  • a subtitle editor
  • a subtitle downloader
  • an open source video player library
  • Video and audio tools
  • Media automation scripts
  • Internationalization libraries
  • OCR screen capture
  • a command line tool for downloading videos