For an open source text based video editor, the strongest matches are mli/autocut (Autocut is a text-based video editor and automatic speech), m-bain/whisperx (WhisperX is an automated speech recognition and alignment toolkit) and ahmetoner/whisper-asr-webservice (This repository provides a self-hosted speech recognition and transcription). nvidia/nemo and funaudiollm/sensevoice round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
We curate open-source GitHub repositories matching “open source alternatives to descript”. Results are ranked by relevance to your query — pick filters below to narrow, or refine with AI.
Autocut is a text-based video editor and automatic speech recognition tool. It allows users to cut and merge video clips by modifying a text transcript instead of using a traditional timeline. The system operates as an FFmpeg video processor and subtitle manipulation utility. It converts spoken audio into text and compacts subtitle files into simplified formats, enabling the removal of unwanted video segments by deleting corresponding sentences from a transcription file. The project covers automated video transcription, non-linear video cutting, and subtitle file management. It supports hard
Autocut is a text-based video editor and automatic speech recognition tool that allows you to edit video by modifying its transcript, though it relies on script-driven trimming rather than a traditional visual timeline interface.
WhisperX is an automated speech recognition toolkit designed to convert spoken audio into text while maintaining precise synchronization with the original media. It functions as an integrated pipeline that combines transcription, phoneme-based alignment, and speaker diarization to produce structured, attributed transcripts. The project distinguishes itself through its use of forced alignment, which matches existing text to audio signals at the phoneme level to generate accurate word-level timestamps. It also incorporates speaker diarization to identify and label unique voices within a recordi
WhisperX is an automated speech recognition and alignment toolkit rather than a full editing application with a timeline and text-based media editing features.
This project provides a self-hosted server for automatic speech recognition, functioning as a containerized inference engine for the Whisper model. It exposes core transcription and translation capabilities through a standardized web interface, allowing for the integration of speech-to-text services into external applications. The service distinguishes itself by incorporating advanced audio analysis tools, including speaker diarization to attribute text to specific individuals and voice activity detection to filter non-speech segments. It supports automated language detection and provides out
This repository provides a self-hosted speech recognition and transcription API rather than a full audio and video editor with a timeline and text-based editing interface.
NeMo is a multimodal AI framework and toolkit designed for the development, training, and scaling of large language models, generative AI systems, and speech-based models. It functions as an automatic speech recognition toolkit, a text-to-speech engine, and a framework for building models that process and generate combinations of text, image, and audio data. The project serves as a conversational AI orchestrator capable of managing real-time, interruptible voice interactions. It provides specialized workflows for speech translation, converting spoken audio from one language into text or speec
This repository is a conversational AI and automatic speech recognition framework rather than a user-facing video and audio editor with transcript-based editing timelines.
SenseVoice is a multilingual speech large language model designed for audio transcription, speaker diarization, and emotion recognition. It functions as an automatic speech recognition system that converts spoken audio into text across multiple languages. The system distinguishes itself by integrating acoustic event detection and speech emotion recognition, allowing it to identify non-speech sounds, such as laughter or applause, and discrete emotional states. It also includes a framework for speaker diarization to track and label different speakers within a single recording. The project's ca
This repository provides a speech recognition and diarization model rather than an end-to-end audio and video editing application with text-based editing timelines.
Whisper-diarization is a system for identifying and separating different speakers in audio recordings by combining OpenAI Whisper for transcription with automated speaker attribution. It functions as a pipeline that isolates vocal tracks from background noise and assigns transcribed segments to specific individuals. The project uses forced alignment to synchronize transcribed text timestamps with audio signals, improving the accuracy of speaker attribution. It employs voice activity detection to separate speech from silence and noise, ensuring precise boundaries for identification. The syste
This project provides speech recognition and speaker diarization pipelines, but it is a library and processing script rather than a complete audio and video editor with timeline editing capabilities.
PaddleSpeech is a comprehensive toolkit of neural models for speech recognition, synthesis, and translation built on the PaddlePaddle deep learning framework. It provides a collection of frameworks and tools for converting spoken audio into written text, synthesizing natural audio from text, and performing direct speech translation. The toolkit includes specialized capabilities for keyword spotting to detect trigger words and speaker verification systems that extract unique voiceprints to identify and distinguish between individuals. It also features end-to-end translation tools that map audi
PaddleSpeech is a speech processing toolkit providing automatic speech recognition and diarization models, but it is a library of neural models rather than a self-contained transcript-based media editor.
whisper.cpp is a C++ implementation of the Whisper speech-to-text model, serving as a lightweight machine learning inference engine and quantized runtime. It provides high-performance automatic speech recognition and real-time audio transcription without requiring a Python environment. The project utilizes model quantization to reduce memory usage and increase inference speed on local hardware. It incorporates hardware acceleration to optimize processing speed across different processors. The system covers audio processing capabilities including voice activity detection, speaker diarization,
Whisper.cpp provides automatic speech recognition and speaker diarization as a low-level inference engine, but it lacks the timeline editor and text-based media editing interface needed for a complete editing suite.
wav2letter is an automatic speech recognition toolkit and deep learning framework designed to convert audio speech signals into written text. It functions as a distributed training system and an inference engine for building and deploying neural network architectures. The system enables the training of large-scale speech models across multiple compute nodes using custom architecture files and structured recipes. It includes an inference engine that allows these trained models to be executed within Python workflows to transform audio sequences into text. The framework covers the full speech r
This repository is a speech recognition framework and toolkit rather than a complete audio and video editor with transcript-based timeline editing.
Auto-subs is an AI transcription and automatic captioning tool that converts spoken audio from video files into synchronized subtitles. It functions as a subtitle generator and a transcription bridge, enabling the conversion of speech to text with automatic speaker identification and multi-language translation support. The software prioritizes data privacy by utilizing on-device AI inference to process audio and video files locally on the user's hardware. It distinguishes itself by offering deep integration with professional video editing workflows, allowing users to export timing and transcr
Auto-subs is a subtitle generator and transcription bridge that integrates with external video editors rather than providing a self-contained timeline-based editor itself, making it a building block for text-driven workflows rather than a complete editing suite.
FunASR is an automatic speech recognition toolkit and multilingual speech-to-text engine designed to convert spoken audio into written text across more than fifty languages. It provides a framework for speaker diarization, an OpenAI-compatible transcription API for local server hosting, and speech models compatible with the ONNX format. The project distinguishes itself by supporting high-performance inference on edge hardware via self-contained binaries and portable model exports. It incorporates specialized capabilities for natural speech generation with adjustable timbre and emotional expre
FunASR provides speech recognition, speaker diarization, and transcription capabilities, but it is an underlying ASR toolkit and engine rather than a complete audio and video editing application with a timeline.
This project is a speech recognition and translation engine that utilizes a sequence-to-sequence transformer architecture to convert audio into text. It is built upon a weakly supervised learning framework, which leverages large-scale, unlabelled audio-transcript data to create generalized speech representations capable of performing simultaneous transcription, language identification, and translation. The system distinguishes itself through a unified multi-task modeling approach that shares token sequences across different objectives, allowing it to handle diverse languages and vocabularies
This repository provides a powerful speech recognition engine for transcribing audio, but it is an AI model and library rather than a complete transcript-based video and audio editing application.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| mli/autocut | 7.6K | Python | apache-2.0 | |
| m-bain/whisperx | 20.2K | Python | bsd-2-clause | |
| ahmetoner/whisper-asr-webservice | 3.3K | Python | MIT | |
| nvidia/nemo | 17.4K | Python | Apache-2.0 | |
| funaudiollm/sensevoice | 7.5K | Python | other | |
| mahmoudashraf97/whisper-diarization | 5.6K | Jupyter Notebook | BSD-2-Clause | |
| paddlepaddle/paddlespeech | 12.6K | Python | Apache-2.0 | |
| ggerganov/whisper.cpp | 50.8K | C++ | MIT | |
| facebookresearch/wav2letter | 6.4K | C++ | NOASSERTION | |
| tmoroney/auto-subs | 2.9K | TypeScript | mit |