For a subtitle editor, the first results are subtitleedit/subtitleedit (Subtitle Edit is a full-featured desktop subtitle editor that supports creating, editing, and synchronizing subtitles with a visual waveform display for alignment, extensive format support, and customization options—exactly what this search asks for), chenyme/chenyme-aavt (Chenyme-AAVT is an AI-powered transcription and translation tool that includes a subtitle formatting editor with real-time video preview, speech recognition for timing, and multilingual support, covering the core needs for subtitle creation and alignment) and aegisub/aegisub (Aegisub is the definitive cross-platform subtitle editor with timeline editing, video preview, multi-format support, style customization, and timing tools — exactly the kind of full-featured captioning tool this search asks for). smacke/ffsubsync and smacke/subsync round out the shortlist. Compare the match explanations and check the project documentation against your requirements.
Open-source software for creating, editing, synchronizing, and formatting subtitle files for various video formats.
Subtitle Edit is a desktop application designed for the creation, synchronization, and adjustment of text-based subtitle files. It provides a graphical interface for managing subtitle workflows, allowing users to modify content and formatting to ensure accurate display during video playback. The application distinguishes itself through a specialized synchronization workflow that utilizes visual waveform displays to align subtitle timestamps with audio and video cues. It supports a wide range of industry-standard file formats, enabling users to convert subtitle data to ensure compatibility acr
Subtitle Edit is a full-featured desktop subtitle editor that supports creating, editing, and synchronizing subtitles with a visual waveform display for alignment, extensive format support, and customization options—exactly what this search asks for.
Chenyme-AAVT is an AI-powered video transcription tool and translation platform. It converts speech from media files into editable text transcripts using speech recognition models and voice activity detection to ensure accurate phrasing and timing. The system functions as a content generator that transforms video transcripts into structured blog posts and marketing graphics using large language models. It also includes a subtitle formatting editor that allows for the modification of subtitle styles with a real-time video preview. The platform provides multilingual translation capabilities th
Chenyme-AAVT is an AI-powered transcription and translation tool that includes a subtitle formatting editor with real-time video preview, speech recognition for timing, and multilingual support, covering the core needs for subtitle creation and alignment.
Cross-platform advanced subtitle editor
Aegisub is the definitive cross-platform subtitle editor with timeline editing, video preview, multi-format support, style customization, and timing tools — exactly the kind of full-featured captioning tool this search asks for.
ffsubsync is a subtitle synchronization tool that aligns subtitle timestamps to audio tracks or reference files using voice activity detection and FFmpeg. It functions as an audio-based subtitle aligner that analyzes speech patterns within a video audio stream to correct timing. The system provides capabilities for cross-language subtitle synchronization, allowing an unsynchronized file to be aligned using a correctly timed subtitle file in a different language as a reference. It also includes a remote media timing engine that streams audio references from network URLs to perform synchronizat
ffsubsync automates subtitle-to-audio synchronization using voice activity detection, but it is a narrow alignment tool rather than a full subtitle editor with timeline editing, video preview, and style customization.
Subsync is a subtitle synchronization tool that aligns subtitle timing to video audio tracks or other synchronized subtitle files. It functions as an audio-based aligner and timing validator to ensure dialogue and captions match during playback. The system utilizes audio-text cross-correlation to match voice activity peaks in audio tracks against subtitle timestamps. It includes a remote media sync client that retrieves files from external servers using standard network protocols for local processing. To ensure accuracy, the tool calculates confidence scores to block updates that fall below
Subsync is a specialized tool for aligning existing subtitle timing to audio tracks, but it does not offer creation, editing, timeline preview, or styling features for a full subtitle editor.
Pyvideotrans is an automated video localization platform designed to transcribe, translate, and dub media content for international distribution. It functions as an end-to-end workflow that combines speech recognition, text translation, and synthetic voice generation to process video files into localized versions. The system distinguishes itself by offering a choice between local model inference for privacy and integration with third-party cloud services via user-provided credentials. This architecture allows users to maintain control over their billing and data security while utilizing modul
Pyvideotrans is an automated pipeline for transcribing, translating, and dubbing videos, which can generate and synchronize subtitles, but it is not a hands-on interactive subtitle editor with timeline, video preview, or style customization.
Whisper-diarization is a system for identifying and separating different speakers in audio recordings by combining OpenAI Whisper for transcription with automated speaker attribution. It functions as a pipeline that isolates vocal tracks from background noise and assigns transcribed segments to specific individuals. The project uses forced alignment to synchronize transcribed text timestamps with audio signals, improving the accuracy of speaker attribution. It employs voice activity detection to separate speech from silence and noise, ensuring precise boundaries for identification. The syste
This is a speaker diarization and transcription pipeline, not a subtitle editor — it lacks the timeline-based editing, video preview, and multi-format subtitle management you need for creating and synchronising subtitles.
WhisperX is an automated speech recognition toolkit designed to convert spoken audio into text while maintaining precise synchronization with the original media. It functions as an integrated pipeline that combines transcription, phoneme-based alignment, and speaker diarization to produce structured, attributed transcripts. The project distinguishes itself through its use of forced alignment, which matches existing text to audio signals at the phoneme level to generate accurate word-level timestamps. It also incorporates speaker diarization to identify and label unique voices within a recordi
WhisperX is an automated speech recognition and forced-alignment toolkit that can generate text with precise timestamps for subtitles, but it is not itself a subtitle editing application—it lacks timeline-based editing, video playback preview, subtitle style customization, and direct export to common subtitle formats, making it a building block rather than the full tool this search is after.
This is a Windows application for automatic speech recognition that transcribes spoken audio from video files into timestamped SRT subtitle files. It serves as a subtitle generator and translation tool that converts media speech into synchronized text. The software functions as a batch media transcriber, allowing the simultaneous processing of multiple audio and video files to generate subtitles in bulk. It includes a translation workflow to convert transcriptions between different languages for the creation of bilingual or localized files. The system also provides text refinement capabiliti
This tool generates SRT subtitles from video audio via speech recognition and supports translation, but it is primarily a transcriber/generator rather than a full subtitle editor with timeline-based editing, manual synchronization, or style customization.
youtube-transcript-api is a Python library designed to retrieve and download subtitles and captions from YouTube videos using video IDs. It functions as an API client that extracts text and timing data for video content. The project includes a wrapper for automated translation, allowing transcripts to be converted into different target languages. It also features a retrieval system that supports routing requests through HTTP, HTTPS, or SOCKS proxies to avoid IP blocking and regional restrictions. The library provides tools for identifying available subtitle tracks and converting raw transcri
This Python library retrieves and translates YouTube captions, but it is a building block for fetching transcripts, not a standalone subtitle editor with timeline editing, video preview, or synchronization features.
Autocut is a text-based video editor and automatic speech recognition tool. It allows users to cut and merge video clips by modifying a text transcript instead of using a traditional timeline. The system operates as an FFmpeg video processor and subtitle manipulation utility. It converts spoken audio into text and compacts subtitle files into simplified formats, enabling the removal of unwanted video segments by deleting corresponding sentences from a transcription file. The project covers automated video transcription, non-linear video cutting, and subtitle file management. It supports hard
Autocut is a text-based video editor that uses automatic speech recognition to generate and edit subtitles as part of its video cutting workflow, but it lacks dedicated timeline-based editing, comprehensive multi-format subtitle support, and styling features that a standalone subtitle editor would offer.
Sewise-player is a comprehensive media playback solution designed for cross-platform video delivery and live broadcast management within web environments. It provides a unified framework for rendering audio and video content, automatically selecting optimal playback technologies to ensure compatibility across diverse browser and device configurations. The player distinguishes itself through a robust programmatic interface that allows developers to manipulate playback states, manage stream switching, and build custom interactive plugins. It supports extensive interface customization, enabling
This is a web-based media player with subtitle display capabilities, not a subtitle creation or editing tool — it plays videos with subtitles but lacks the timeline editing, alignment, and authoring features you need.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| subtitleedit/subtitleedit | 13.2K | C# | MIT | |
| chenyme/chenyme-aavt | 2.9K | Python | mit | |
| 3.4K |
| C++ |
| NOASSERTION |
| smacke/ffsubsync | 7.6K | Python | mit |
| smacke/subsync | 7.7K | Python | MIT |
| jianchang512/pyvideotrans | 18K | Python | GPL-3.0 |
| mahmoudashraf97/whisper-diarization | 5.6K | Jupyter Notebook | BSD-2-Clause |
| m-bain/whisperx | 20.2K | Python | bsd-2-clause |
| wxbool/video-srt-windows | 5K | Go | GPL-2.0 |
| jdepoix/youtube-transcript-api | 6.9K | Python | mit |