awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
WEIFENG2333 avatar

WEIFENG2333/VideoCaptioner

0
View on GitHub↗
13,278 Stars·1,081 Forks·Python·gpl-3.0·13 Aufrufewww.videocaptioner.cn↗

VideoCaptioner

VideoCaptioner is an automated tool designed to generate and embed time-synchronized subtitles into video files. By leveraging speech recognition models, the software converts spoken audio into text and calculates precise timestamps to ensure captions align with the original media.

The project operates as a local-first inference pipeline, performing all transcription tasks on the host machine to maintain data privacy. It utilizes a transformer-based neural network for speech recognition and integrates a multimedia framework to handle the technical aspects of video processing and subtitle stream multiplexing.

Beyond automated transcription, the tool provides capabilities for hardcoded subtitle embedding and the permanent integration of text tracks into video containers. This functionality ensures that generated captions remain visible across various media players and devices, supporting accessibility for hearing-impaired viewers.

Features

  • Automated Subtitle Generators - Uses speech recognition models to transcribe audio and embed time-synced captions directly into video files.
  • Audio and Video Processors - Provides media manipulation capabilities to merge subtitle tracks into video containers for permanent caption visibility.
  • Whisper-Based Engines - Converts spoken audio into text using advanced machine learning models for accurate subtitle generation.
  • Automated Video Transcribers - Converts spoken audio from video files into accurate, time-synced text files using automated speech recognition.
  • Timestamped Subtitle Generators - Transcribes audio from video files using automated speech recognition to produce accurate, time-synced subtitle files.
  • Audio and Subtitle Tools - Video subtitle processing assistant powered by large language models.
  • Hardcoded Embedders - Merges subtitle tracks directly into video files to ensure captions remain permanently visible across any media player.
  • Subtitle Processing - Merges subtitle tracks directly into video files to ensure captions remain permanently visible across any media player or device.
  • Synchronization Engines - Calculates precise start and end timestamps to ensure transcribed text remains perfectly synchronized with video playback.
  • Video Accessibility Tools - Enhances video accessibility for hearing-impaired viewers by generating and permanently attaching text captions to media files.
  • Multimedia Processing Suites - Provides command-line multimedia processing suites to handle complex video transcoding and stream manipulation tasks.

Star-Verlauf

Star-Verlauf für weifeng2333/videocaptionerStar-Verlauf für weifeng2333/videocaptioner

KI-Suche

Entdecke weitere awesome Repositories

Beschreibe in einfachen Worten, was du brauchst — die KI bewertet tausende kuratierte Open-Source-Projekte nach Relevanz.

Start searching with AI

Open-Source-Alternativen zu VideoCaptioner

Ähnliche Open-Source-Projekte, sortiert nach der Anzahl der gemeinsamen Funktionen mit VideoCaptioner.
  • wxbool/video-srt-windowsAvatar von wxbool

    wxbool/video-srt-windows

    5,037Auf GitHub ansehen↗

    This is a Windows application for automatic speech recognition that transcribes spoken audio from video files into timestamped SRT subtitle files. It serves as a subtitle generator and translation tool that converts media speech into synchronized text. The software functions as a batch media transcriber, allowing the simultaneous processing of multiple audio and video files to generate subtitles in bulk. It includes a translation workflow to convert transcriptions between different languages for the creation of bilingual or localized files. The system also provides text refinement capabiliti

    Goffmpeggogolang
    Auf GitHub ansehen↗5,037
  • linyqh/narratoaiAvatar von linyqh

    linyqh/NarratoAI

    8,091Auf GitHub ansehen↗

    NarratoAI is an automated video production pipeline that uses large language models to generate scripts, voiceovers, and edited video commentary. It functions as a combined scriptwriter, voiceover generator, and video editor to streamline the creation of movie and television commentary content. The system automates the production workflow by converting input data into structured narrative scripts, synthesizing artificial speech for narration, and programmatically assembling video clips based on script timestamps. It also converts spoken audio from video files into written text for subtitles a

    Pythonaiagentaiopsgemini-api
    Auf GitHub ansehen↗8,091
  • buxuku/smartsubAvatar von buxuku

    buxuku/SmartSub

    4,056Auf GitHub ansehen↗

    SmartSub is a cross-platform desktop application for AI-driven video transcription and subtitle generation. It converts audio and video files into text subtitles using local AI models and incorporates hardware acceleration to increase processing speed. The tool features a subtitle translator that leverages large language models, such as OpenAI and DeepSeek, to convert subtitles between different languages. It includes a visual editor for proofreading and polishing transcribed text, paired with a video preview for frame-accurate synchronization. The software supports batch processing of multi

    TypeScriptdeepseekelectronnodejs
    Auf GitHub ansehen↗4,056
  • agermanidis/autosubAvatar von agermanidis

    agermanidis/autosub

    4,197Auf GitHub ansehen↗

    Autosub is a command-line media processor and automatic subtitle generator that converts audio streams from video and audio files into timed text overlays. It functions as an AI speech-to-text converter that uses OpenAI Whisper to generate synchronized subtitles. The tool includes a language translation pipeline to convert transcribed speech into target languages, enabling multilingual video captioning. It manages the process from audio-stream extraction to the serialization of final subtitle files for local storage. The system covers audio-to-text transcription, time-stamped text mapping, a

    Python
    Auf GitHub ansehen↗4,197
Alle 30 Alternativen zu VideoCaptioner anzeigen→

Häufig gestellte Fragen

Was macht weifeng2333/videocaptioner?

VideoCaptioner is an automated tool designed to generate and embed time-synchronized subtitles into video files. By leveraging speech recognition models, the software converts spoken audio into text and calculates precise timestamps to ensure captions align with the original media.

Was sind die Hauptfunktionen von weifeng2333/videocaptioner?

Die Hauptfunktionen von weifeng2333/videocaptioner sind: Automated Subtitle Generators, Audio and Video Processors, Whisper-Based Engines, Automated Video Transcribers, Timestamped Subtitle Generators, Audio and Subtitle Tools, Hardcoded Embedders, Subtitle Processing.

Welche Open-Source-Alternativen gibt es zu weifeng2333/videocaptioner?

Open-Source-Alternativen zu weifeng2333/videocaptioner sind unter anderem: wxbool/video-srt-windows — This is a Windows application for automatic speech recognition that transcribes spoken audio from video files into… linyqh/narratoai — NarratoAI is an automated video production pipeline that uses large language models to generate scripts, voiceovers,… buxuku/smartsub — SmartSub is a cross-platform desktop application for AI-driven video transcription and subtitle generation. It… agermanidis/autosub — Autosub is a command-line media processor and automatic subtitle generator that converts audio streams from video and… browser-use/video-use — This project is an AI video post-production suite that uses large language models and programmatic tools to automate… systran/faster-whisper — Faster-Whisper is a high-performance implementation of the Whisper speech-to-text model designed for efficient audio…