14 Repos
Automated workflows for performing the same operations across multiple video files.
Distinct from Batch Processing: Existing batch candidates are for audio, cloud trials, or data inputs, not video file sets.
Explore 14 awesome GitHub repositories matching graphics & multimedia · Batch Video Processing. Refine with filters or upvote what's useful.
Perfect Green Screen Keys
Processes multiple video files in one run, applying background removal to each clip automatically.
This project is an optical character recognition tool designed to extract hardcoded subtitles from video frames and convert them into synchronized subtitle files. It functions as a text processor that transforms embedded visual text into a written format to improve video accessibility and translation. The system uses graphics processing units to increase the speed and accuracy of text recognition. It includes a subtitle cleaning tool that applies custom mapping configurations to filter out watermarks, channel logos, and duplicate lines from the extracted text. The tool supports batch process
Enables subtitles to be extracted from multiple video files simultaneously when resolution and text regions are identical.
Gyroflow is a gyroscope video stabilization software and IMU telemetry processor designed to remove camera shake from video files. It functions as a hardware-accelerated video renderer and lens calibration tool, utilizing embedded or external gyroscope and accelerometer data to perform pixel-level stabilization. The system is distinguished by its ability to integrate with professional non-linear video editing software via plugins, allowing stabilization to be applied directly to timelines without transcoding original footage. It supports diverse telemetry ingestion from camera brands, flight
Applies uniform smoothness, horizon locking, and zoom configurations across multiple video files.
Backgroundremover is an AI-powered tool that removes backgrounds from both images and videos, accessible through a command-line interface and a Python API. At its core, it uses a pre-trained deep learning model to classify each pixel as foreground or background, producing a binary mask for removal. The tool distinguishes itself through multiple integration methods and output capabilities. It can process images and videos via Unix pipeline data streams, operate as an HTTP API server, or be called programmatically within Python scripts. Users can choose among different AI models to balance proc
Processes images and videos from the terminal with batch operations, Unix pipes, and model selection.
MoneyPrinterPlus is an automated video production system designed for the mass creation of short-form AI content. It functions as an end-to-end pipeline that uses large language models to generate scripts, synthesize voiceovers, and produce visual assets to assemble complete videos. The project is distinguished by its ability to batch-process high volumes of unique content through automated mixing and randomized asset pairing. It includes a social media auto-publisher that uses browser simulation to automate the upload and distribution of generated videos to platforms such as TikTok and Xiaoh
Produces large quantities of non-duplicate short videos through automated mixing and batch editing.
Manages multiple encoding jobs in a queue for efficient batch processing of video files.
Dies ist eine Windows-Anwendung für automatische Spracherkennung, die gesprochenes Audio aus Videodateien in zeitgestempelte SRT-Untertiteldateien transkribiert. Sie dient als Untertitelgenerator und Übersetzungstool, das Medien-Sprache in synchronisierten Text umwandelt. Die Software fungiert als Batch-Medien-Transkribierer, der die gleichzeitige Verarbeitung mehrerer Audio- und Videodateien ermöglicht, um Untertitel in großen Mengen zu generieren. Sie enthält einen Übersetzungsworkflow zur Konvertierung von Transkriptionen zwischen verschiedenen Sprachen für die Erstellung zweisprachiger oder lokalisierter Dateien. Das System bietet zudem Textverfeinerungsfunktionen unter Verwendung regulärer Ausdrücke und benutzerdefinierter Filter, um Transkripte durch das Entfernen von Füllwörtern und unerwünschten Mustern zu bereinigen. Dies wird durch eine native grafische Windows-Benutzeroberfläche unterstützt.
Automates the creation of subtitles across multiple video files using batch workflows.
node-ytdl-core ist eine JavaScript-Bibliothek für Node.js, die entwickelt wurde, um Metadaten zu extrahieren und Video- sowie Audioinhalte von YouTube zu streamen. Sie dient als Media-Downloader und Stream-Fetcher und ermöglicht es Nutzern, Videodetails und Mediendaten aus Remote-Quellen abzurufen. Die Bibliothek bietet spezialisierte Funktionen für die Videoextraktion, einschließlich der Fähigkeit, Medien-URLs nach eindeutigen Identifikatoren zu parsen und verfügbare Formate zu analysieren. Sie ermöglicht die Auswahl und Filterung spezifischer Video- und Audiostreams basierend auf Qualitäts- und Auflösungskriterien. Das Projekt verwaltet den Netzwerkverkehr durch die Vermeidung von Rate-Limits und die Verwaltung von Authentifizierungs-Cookies. Es nutzt lesbare Streams für das Data-Piping und unterstützt Byte-Range-Requests, um spezifische Segmente von Mediendateien abzurufen.
Parses URLs and strings to retrieve one or more unique video identifiers for batch retrieval.
Automatic Optical Disc Ripping Server is a headless system that detects inserted CDs, DVDs, and Blu-rays to automatically extract media, transcode video, and eject discs. It functions as a multi-drive media digitizer using a concurrent processing pipeline to rip and transcode media from several optical drives simultaneously without queuing. The system includes an asynchronous video transcoding pipeline that batches conversion tasks to run during scheduled off-peak hours. It also serves as a media server automation tool, fetching metadata from online APIs to name folders and trigger library re
Implements a batch processing system that converts ripped video files into target formats during scheduled off-peak hours.
SmartSub ist eine plattformübergreifende Desktop-Anwendung für KI-gestützte Videotranskription und Untertitelgenerierung. Sie konvertiert Audio- und Videodateien mithilfe lokaler KI-Modelle in Text-Untertitel und nutzt Hardwarebeschleunigung, um die Verarbeitungsgeschwindigkeit zu erhöhen. Das Tool verfügt über einen Untertitel-Übersetzer, der Large Language Models wie OpenAI und DeepSeek nutzt, um Untertitel zwischen verschiedenen Sprachen zu konvertieren. Es enthält einen visuellen Editor zum Korrekturlesen und Verfeinern transkribierter Texte, gepaart mit einer Videovorschau für frame-genaue Synchronisation. Die Software unterstützt die Stapelverarbeitung mehrerer Mediendateien und bietet Utilities zum Einbetten von Untertiteln als schaltbare Soft-Tracks oder zum festen Einbrennen in Videoframes. Zudem enthält sie Systeme zur Verwaltung von Transkriptionsmodelldateien und zur Konfiguration externer KI-Service-Parameter.
Automates the transcription and translation process across multiple video files simultaneously.
Squirrel-RIFE is a GPU-accelerated video processing tool that uses a neural network to generate intermediate frames between existing video frames, enabling smooth slow-motion effects and frame rate conversion. It is built around the RIFE (Real-Time Intermediate Flow Estimation) model, which analyzes motion between consecutive frames to predict and insert new frames, and leverages NVIDIA CUDA for parallel processing to achieve high-speed inference. The tool distinguishes itself by combining neural frame interpolation with practical video preprocessing features, including pixel-level duplicate
Processes multiple video frames in parallel batches on NVIDIA GPUs to maximize throughput and reduce per-frame overhead.
Short video factory is a local AI content generator and automated video editing tool. It provides a production pipeline that uses large language models to transform text prompts into marketing scripts and rendered short-form videos. The system is designed for local-first execution, running all processing and asset management on the host machine to maintain data privacy. It distinguishes itself through a batch-processing workflow that can sequentially execute copywriting and rendering for multiple items using predefined presets. The software covers a broad range of media capabilities, includi
Automatically produces a sequence of videos by batching the copywriting and rendering processes.
ComfyUI-SeedVR2_VideoUpscaler is an AI video upscaling tool that uses diffusion models to increase the resolution of videos and images while maintaining visual consistency across frames. The project implements distributed video rendering by splitting datasets into chunks for parallel processing across multiple GPUs. It utilizes model compilation and specialized attention backends to reduce inference latency and increase throughput. Additional capabilities include video color correction using wavelet and LAB matching methods to preserve color fidelity. Hardware memory is managed through block
Offers automated workflows for performing resolution enhancement across multiple video files via CLI.
Dieses Projekt ist ein Deep-Learning-Framework für die Erkennung und Verfolgung menschlicher Körper-Keypoints in Bildern und Videostreams. Es fungiert sowohl als Echtzeit-Motion-Tracking-System als auch als Machine-Learning-Umgebung für das Training und die Evaluierung von Pose-Estimation-Modellen. Das System nutzt ein Convolutional Neural Network mit zwei Zweigen, um Körperteilpositionen und deren gerichtete Verbindungen gleichzeitig vorherzusagen. Es verwendet mehrstufige Feature-Verfeinerung, um die Genauigkeit der Keypoint-Lokalisierung zu verbessern, und nutzt Greedy-Parsing- sowie Bipartite-Matching-Algorithmen, um erkannte Teile zu individuellen Skeletten zuzuordnen. Um die Leistung während der Live-Videoanalyse aufrechtzuerhalten, führt das Framework parallele Inferenz über Bildregionen hinweg mittels Tensor-basierter Batch-Verarbeitung aus. Über das Echtzeit-Tracking hinaus bietet die Bibliothek Tools für das Training von Modellen auf annotierten Datensätzen und die Berechnung der Mean Average Precision (mAP) gegenüber standardisierten Benchmarks, um die Erkennungsqualität zu verifizieren. Das Repository enthält die notwendigen Komponenten, um den gesamten Lebenszyklus der Pose Estimation zu verwalten, vom anfänglichen Modelltraining bis zur Leistungsvalidierung.
Executes parallel tensor-based batch processing to maintain high frame rates during real-time video analysis.