30 open-source projects similar to irt-open-source/scf, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.
() | | / _| | | | ( | | | ' ` \ / | / | | | \ \ | | | | | | | | \ \ | ( | || | ) | || || || || |/ \| \/ |/
$$\ $$\ $$ | $$ | $$$$$$\ $$$$$$\ $$$$$$$\ $$$$$$\ $$$$$$$\ $$\ $$\ \$$ |\$$ | $$ |$$ $$\ $$ $$\\$$\ $$ | $$ | $$ | $$ / $$ / $$ |$$ | $$ |\$$\$$ / $$ |$$\ $$ |$$\ $$ | $$ | $$ |$$ | $$ | \$$$ / \$$$$ |\$$$$ |\$$$$$$$\ \$$$$$$ |$$ | $$ | \$ / \/ \/ \_| \/ \| \| \/
Please note: This code was incorporated into the GStreamer gst-plugins-bad plugin set in 2016 and, therefore, this repository is not in active development
This is the repository for the interoperability kit of EBU-TT Live.
The IRT-EBU-TT-D Application Samples are published to support early EBU-TT-D implementation by service developers.
A Web UI for easy subtitle using whisper model.
Pyvideotrans is an automated video localization platform designed to transcribe, translate, and dub media content for international distribution. It functions as an end-to-end workflow that combines speech recognition, text translation, and synthetic voice generation to process video files into localized versions. The system distinguishes itself by offering a choice between local model inference for privacy and integration with third-party cloud services via user-provided credentials. This architecture allows users to maintain control over their billing and data security while utilizing modul
Transcription, forced alignment, and audio indexing with OpenAI's Whisper
WhisperX is an automated speech recognition toolkit designed to convert spoken audio into text while maintaining precise synchronization with the original media. It functions as an integrated pipeline that combines transcription, phoneme-based alignment, and speaker diarization to produce structured, attributed transcripts. The project distinguishes itself through its use of forced alignment, which matches existing text to audio signals at the phoneme level to generate accurate word-level timestamps. It also incorporates speaker diarization to identify and label unique voices within a recordi
A collection of related tools that provide support for or make use of the W3C Timed Text Markup Language (TTML).
ffsubsync is a subtitle synchronization tool that aligns subtitle timestamps to audio tracks or reference files using voice activity detection and FFmpeg. It functions as an audio-based subtitle aligner that analyzes speech patterns within a video audio stream to correct timing. The system provides capabilities for cross-language subtitle synchronization, allowing an unsynchronized file to be aligned using a correctly timed subtitle file in a different language as a reference. It also includes a remote media timing engine that streams audio references from network URLs to perform synchronizat