awesome-repositories.com
المدونة
MCP
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعخادم MCPحولكيفية ترتيب النتائجالصحافة
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

26 مستودعات

Awesome GitHub RepositoriesVideo Processing Tools

Specialized software tools for navigating, editing, and analyzing individual frames within video files.

Explore 26 awesome GitHub repositories matching graphics & multimedia · Video Processing Tools. Refine with filters or upvote what's useful.

Awesome Video Processing Tools GitHub Repositories

اعثر على أفضل المستودعات باستخدام الذكاء الاصطناعي.سنبحث عن أفضل المستودعات المطابقة باستخدام الذكاء الاصطناعي.
  • ffmpeg/ffmpegالصورة الرمزية لـ FFmpeg

    FFmpeg/FFmpeg

    61,176عرض على GitHub↗

    FFmpeg is a cross-platform multimedia framework designed for the recording, conversion, and streaming of audio and video content. It functions as a comprehensive toolkit that provides both a command-line utility for direct media manipulation and a collection of low-level libraries for integration into custom applications. At its core, the project utilizes a packet-based stream engine and a format-agnostic abstraction layer to handle diverse media standards, containers, and network protocols. The framework distinguishes itself through a modular, graph-based filter execution model that allows f

    Processes video frames through a chain of filters to perform visual transformations, effects, and adjustments.

    Caudiocffmpeg
    عرض على GitHub↗61,176
  • deepfakes/faceswapالصورة الرمزية لـ deepfakes

    deepfakes/faceswap

    55,289عرض على GitHub↗

    Faceswap is a comprehensive framework for automated media manipulation and neural face synthesis. It provides a modular pipeline that manages the entire lifecycle of facial feature extraction, deep learning model training, and image conversion. By coordinating complex computer vision workflows, the system enables users to map facial identities between source and destination datasets while maintaining structural alignment and lighting consistency across video frames. The project distinguishes itself through a highly extensible plugin-based architecture that handles hardware-accelerated process

    Navigates video frames via playback controls and filtering to pinpoint specific segments for manual editing.

    Pythondeep-face-swapdeep-learningdeep-neural-networks
    عرض على GitHub↗55,289
  • leandromoreira/digital_video_introductionالصورة الرمزية لـ leandromoreira

    leandromoreira/digital_video_introduction

    16,232عرض على GitHub↗

    This project is an educational suite and technical guide designed for mastering video codecs and signal processing. It provides a structured curriculum through an engineering course, interactive labs, and tutorials focused on the fundamental principles of video compression and digital signal processing. The resource includes a technical guide for analyzing specific codecs like AV1, VP9, and H.265. It distinguishes itself by providing a containerized media lab, which ensures a consistent development environment for experimenting with video technology tools and notebooks. The project covers a

    Provides command-line tools to cut, download, and encode video clips to modify media properties.

    Jupyter Notebookadaptive-streamingarithmetic-codingaudio
    عرض على GitHub↗16,232
  • cvat-ai/cvatالصورة الرمزية لـ cvat-ai

    cvat-ai/cvat

    15,317عرض على GitHub↗

    CVAT is an open-source, web-based platform designed for annotating images, videos, and 3D point clouds to create high-quality training datasets for machine learning. It functions as a containerized server that orchestrates the entire lifecycle of computer vision data, from initial task creation and manual labeling to quality assurance and final dataset export. The platform distinguishes itself through deep integration with machine learning models, allowing users to deploy custom AI models as serverless functions for automated object detection, tracking, and skeleton annotation. It supports co

    Provides precise controls for stepping through, seeking, and inspecting individual frames within video sequences.

    Pythonannotationannotation-toolannotations
    عرض على GitHub↗15,317
  • zulko/moviepyالصورة الرمزية لـ Zulko

    Zulko/moviepy

    14,699عرض على GitHub↗

    MoviePy is a Python video editing library and automated video processor designed for programmatically cutting, concatenating, and manipulating video and audio files. It serves as a non-linear video editor and an interface for FFmpeg to handle the reading, writing, and conversion of diverse media formats and codecs. The library enables automated video composition through the layering of multiple video and audio streams using transparency and coordinate-based positioning. It supports dynamic content generation by inserting text overlays and performing custom video frame processing where raw fra

    Represents video frames as NumPy arrays to enable high-performance, pixel-level transformations and custom effects.

    Pythonanimationgifhacktoberfest
    عرض على GitHub↗14,699
  • chocobozzz/peertubeالصورة الرمزية لـ Chocobozzz

    Chocobozzz/PeerTube

    14,520عرض على GitHub↗

    PeerTube is a decentralized, open-source video hosting platform that enables users to operate independent, interoperable servers. By utilizing the ActivityPub protocol, it connects these servers into a global, federated network where users can follow channels, discover content, and interact across different instances. The platform is designed to function as a self-hosted video content management system, providing a community-driven alternative to centralized media services. What distinguishes PeerTube is its hybrid approach to content delivery and infrastructure management. It integrates peer

    Provides a media player with playback controls, keyboard shortcuts, and visual navigation tools.

    TypeScriptactivitypubangulardecentralized
    عرض على GitHub↗14,520
  • leandromoreira/ffmpeg-libav-tutorialالصورة الرمزية لـ leandromoreira

    leandromoreira/ffmpeg-libav-tutorial

    11,011عرض على GitHub↗

    هذا المشروع عبارة عن دليل هندسة وسائط قائم على C وإطار عمل لمعالجة الوسائط المتعددة مصمم لإدارة برامج الترميز (codecs)، والإطارات، والحزم داخل نظام FFmpeg وLibav. يوفر وثائق تقنية وأنماط تنفيذ لتحويل الترميز (transcoding)، وإعادة التغليف (remuxing)، وتغيير حجم بيانات الفيديو والصوت. يتضمن المشروع بيئة تطوير بالحاويات تغلف مكتبات الوسائط وسلاسل الأدوات المطلوبة داخل صورة افتراضية لضمان بيئات بناء متسقة. يغطي إطار العمل مجموعة من تدفقات عمل هندسة الوسائط المتعددة، بما في ذلك البث بمعدل بت تكيفي، وإعادة تغليف حاويات الوسائط، واستخراج البيانات الوصفية. كما يوفر قدرات لمعالجة التدفق، مثل مزامنة الصوت والفيديو وتغيير حجم الدقة، إلى جانب أدوات لتسجيل تفاصيل برامج الترميز والتوقيت.

    Adjusts the spatial resolution of video frames to fit specific display sizes or bandwidth limits.

    C
    عرض على GitHub↗11,011
  • kkroening/ffmpeg-pythonالصورة الرمزية لـ kkroening

    kkroening/ffmpeg-python

    10,999عرض على GitHub↗

    ffmpeg-python is a Python wrapper that translates programmatic method calls into command-line arguments for executing FFmpeg media processing tasks. It functions as a multimedia transcoding interface and a media stream capture tool, allowing for the recording of live audio and video from hardware devices and network sources. The library features a fluent interface for constructing complex directed graphs of audio and video filters through method chaining. It also includes an FFprobe metadata extractor that retrieves structured technical properties from media files and returns them as Python d

    Extracts individual frames of raw video and audio data into numerical arrays for external analysis.

    Python
    عرض على GitHub↗10,999
  • lengstrom/fast-style-transferالصورة الرمزية لـ lengstrom

    lengstrom/fast-style-transfer

    10,963عرض على GitHub↗

    This project is a TensorFlow-based neural style transfer framework designed to apply the artistic textures and colors of a painting to images and videos. It utilizes a feed-forward image stylizer that transforms visual appearance in a single pass, avoiding the need for iterative optimization. The system includes a deep learning training pipeline that teaches convolutional neural networks to replicate specific styles using perceptual loss functions. It also features a video frame processor that decomposes video files into individual images for sequential stylization and reassembly. The softwa

    Implements a processing loop that applies artistic style filters to a sequence of video frames.

    Pythondeep-learningneural-networksneural-style
    عرض على GitHub↗10,963
  • owncast/owncastالصورة الرمزية لـ owncast

    owncast/owncast

    10,950عرض على GitHub↗

    Owncast is a self-hosted live streaming server that provides full control over broadcast infrastructure and audience data. It functions as an RTMP video streaming server, accepting incoming video feeds and distributing them to viewers through HLS-based segmented streaming. The platform includes a built-in, stateful web-based chat interface that enables real-time viewer engagement during broadcasts. The project distinguishes itself through deep integration with the decentralized Fediverse, allowing servers to automatically broadcast stream status updates and notify followers across distributed

    Delegates resource-intensive video transcoding to external hardware or software to preserve local system performance.

    Goactivitypubbroadcastingchat
    عرض على GitHub↗10,950
  • eduardolundgren/tracking.jsالصورة الرمزية لـ eduardolundgren

    eduardolundgren/tracking.js

    9,472عرض على GitHub↗

    tracking.js is a browser computer vision library written in JavaScript for performing real-time image analysis and object tracking directly within a web browser. It functions as a real-time object tracker, a color tracking tool, and a face detection utility. The library enables the detection and monitoring of specific color ranges, human faces, and known visual patterns across consecutive video frames. It extracts visual features and descriptors from images to identify distinct landmarks for matching and tracking. The project covers broad computer vision capabilities, including the ability t

    Implements a frame-by-frame processing loop for real-time video stream analysis.

    JavaScript
    عرض على GitHub↗9,472
  • peterl1n/robustvideomattingالصورة الرمزية لـ PeterL1n

    PeterL1n/RobustVideoMatting

    9,244عرض على GitHub↗

    RobustVideoMatting is a deep learning video matting tool and PyTorch library designed to remove backgrounds from videos and extract human subjects. It utilizes a temporal video segmentation model to ensure consistent matting and reduce flickering across video frames. The project includes a cross-platform model exporter that converts trained neural networks into various runtime formats. This allows for model deployment across multiple environments, including web and mobile applications. The framework provides capabilities for temporal video background removal and AI video post-production with

    Processes sequential video frames while maintaining state to ensure temporal consistency across the video.

    Pythonaicomputer-visiondeep-learning
    عرض على GitHub↗9,244
  • shopify/react-native-skiaالصورة الرمزية لـ Shopify

    Shopify/react-native-skia

    8,424عرض على GitHub↗

    react-native-skia is a cross-platform graphics framework that provides a high-performance 2D graphics engine for rendering shapes, paths, and images. It functions as a vector graphics engine and UI animation toolkit, allowing for hardware-accelerated visuals across mobile and web platforms. The project is distinguished by its integration of the Skia 2D graphics library, enabling a shader and filter pipeline for complex pixel-level effects. It supports the rendering of Lottie animations exported from After Effects and the execution of animations directly on the UI thread to maintain fluid moti

    Resizes drawings between bounding rectangles using contain or cover fit modes.

    TypeScriptreactreact-nativeskia
    عرض على GitHub↗8,424
  • fluent-ffmpeg/node-fluent-ffmpegالصورة الرمزية لـ fluent-ffmpeg

    fluent-ffmpeg/node-fluent-ffmpeg

    8,251عرض على GitHub↗

    node-fluent-ffmpeg هو غلاف Node.js لـ FFmpeg يوفر واجهة سلسة لتنفيذ أوامر الوسائط ومعالجة الملفات. يعمل كمدير عمليات يتعامل مع دورة حياة ثنائيات FFmpeg الخارجية، مما يتيح تحويل الوسائط برمجياً، وتوليد صور مصغرة للفيديو، واستخراج البيانات الوصفية عبر ffprobe. تتميز المكتبة بمنشئ أوامر يترجم استدعاءات أساليب JavaScript إلى وسيطات سطر أوامر. تتميز بمراقبة التقدم القائمة على الأحداث لتتبع الإطارات المعالجة والإنتاجية، بالإضافة إلى القدرة على توجيه بيانات الوسائط المعالجة مباشرة إلى تدفقات قابلة للكتابة للتعامل في الوقت الفعلي. يغطي المشروع قدرات واسعة لمعالجة الوسائط، بما في ذلك إعداد التشفير لخصائص الصوت والفيديو، وتعريفات مخططات التصفية المعقدة للتأثيرات المرئية والصوتية، وإدارة المدخلات لربط مصادر متعددة. يتضمن أيضاً أدوات لفحص حاويات الوسائط والتدفقات لاسترجاع البيانات الوصفية التقنية.

    Provides tools for adjusting the spatial resolution and aspect ratio of video frames.

    JavaScript
    عرض على GitHub↗8,251
  • deniscerri/ytdlnisالصورة الرمزية لـ deniscerri

    deniscerri/ytdlnis

    7,742عرض على GitHub↗

    ytdlnis is a mobile application that serves as a graphical client for the yt-dlp engine on Android. It functions as a media downloader and manager, providing a user interface to retrieve video and audio from websites. The project distinguishes itself by integrating directly with the Android system share menu and intents to trigger background downloads from external apps. It includes a dedicated authentication cookie manager to import and sync browser session data, enabling the retrieval of private, age-restricted, or premium content. The application covers broad capability areas including au

    Provides tools for trimming segments, removing sponsored filler, and embedding subtitles.

    Kotlinandroidaudiodownloader
    عرض على GitHub↗7,742
  • serpentai/serpentaiالصورة الرمزية لـ SerpentAI

    SerpentAI/SerpentAI

    6,979عرض على GitHub↗

    SerpentAI is a game AI development kit and computer vision framework designed for building autonomous agents that interact with video games. It serves as a game input automation tool and a machine learning model integration engine, allowing developers to create agents that perceive game states and execute actions. The framework utilizes a plugin-based agent architecture to provide modular extensions for game-specific logic and behaviors. It features a specialized system for training, bundling, and deploying machine learning classifiers to recognize visual contexts and game states in real time

    Captures screen regions and passes them through a handler loop for real-time computer vision analysis.

    Pythonartificial-intelligencecomputer-visiondeep-learning
    عرض على GitHub↗6,979
  • mifi/editlyالصورة الرمزية لـ mifi

    mifi/editly

    5,435عرض على GitHub↗

    Editly هو محرك فيديو برمجي بدون واجهة رسومية (headless) ومجمع آلي للفيديوهات. يعمل كمحرر فيديو تعريفي (declarative) يقوم بإنشاء ملفات MP4 وGIF من بيانات مهيكلة أو كود، مما يلغي الحاجة إلى واجهة مستخدم رسومية يدوية. يتميز النظام بقدرته على دمج تظليل GLSL (fragment shaders) كطبقات مرئية ضمن جدول زمني برمجي. يستخدم نموذجاً يعتمد على الإعدادات لتعريف المقاطع والطبقات والمسارات الصوتية، مما يسمح بتجميع الفيديوهات بشكل قابل للتكرار وإنشاء رسومات برمجية مخصصة. يغطي المحرك مجموعة واسعة من قدرات إنتاج الوسائط، بما في ذلك تسلسل الفيديو غير الخطي، وتغيير حجم الأصول، ودمج الصوت متعدد المسارات مع التطبيع التلقائي (normalization) وخفض الصوت (ducking). كما يتعامل مع التكوين المرئي من خلال العرض القائم على اللوحة (canvas-based rendering) لتراكبات النصوص، والترجمات، وتأثيرات محاكاة الكاميرا مثل التكبير والتحريك. محرك العرض وتبعياته متاحان كصورة حاوية (containerized image) لضمان التنفيذ المتسق عبر بيئات مختلفة.

    Resizes video frames using various fit modes, such as stretch or cover with background blur, to match target dimensions.

    TypeScript
    عرض على GitHub↗5,435
  • vanilagy/mediabunnyالصورة الرمزية لـ Vanilagy

    Vanilagy/mediabunny

    5,254عرض على GitHub↗

    This is a cross-platform media processing library that reads, writes, encodes, and decodes media in both browser and server environments. It supports common container formats including ISOBMFF, Matroska, Ogg, MPEG-TS, and HLS, and handles codec operations through a combination of WebCodecs API and WebAssembly-based encoders. Media is processed in streaming pipelines that maintain constant memory usage and automatically apply backpressure from output speed to all upstream components. The library distinguishes itself through a plugin-based codec registration system that allows extending support

    FFmpeg.wasm resizes and crops video samples with options for fit mode using a simple method.

    TypeScriptaudiodecodingdemuxing
    عرض على GitHub↗5,254
  • tachibanayoshino/animeganالصورة الرمزية لـ TachibanaYoshino

    TachibanaYoshino/AnimeGAN

    4,603عرض على GitHub↗

    AnimeGAN is a generative adversarial network and image translator developed with TensorFlow. It is designed for photo-to-anime style transfer, utilizing a deep learning system to transform real-world photographs and video frames into anime-style imagery. The system includes a video-to-anime converter that applies consistent visual transformations across sequential frames. It supports both the training of generative networks on artistic datasets to replicate specific styles and the extraction of generator weights from checkpoints for efficient inference. The project provides utilities for ima

    Applies generative transformations across a timeline by processing video files as a sequence of individual frames.

    Pythonanime-imagesanimeganhayao-style
    عرض على GitHub↗4,603
  • microsoft/vottالصورة الرمزية لـ microsoft

    microsoft/VoTT

    4,427عرض على GitHub↗

    VoTT هو برنامج لتعليق بيانات الرؤية الحاسوبية (computer vision annotation) وأداة لإعداد مجموعات بيانات تعلم الآلة. هو تطبيق سطح مكتب مصمم لرسم مربعات الإحاطة (bounding boxes) وتعيين وسوم للكائنات في الصور ومقاطع الفيديو لإنشاء مجموعات بيانات تدريب لنماذج اكتشاف الكائنات. يستخدم التطبيق واجهة سطح مكتب متعددة المنصات لإدارة أصول الصور والفيديو. ويتميز بتكامل تخزين محلي (local-first) للتعامل مع أصول الوسائط الكبيرة مباشرة من نظام ملفات الجهاز المضيف، ويتضمن أخذ عينات فيديو محكوم بمعدل الإطارات لاستخراج صور محددة من تدفقات الفيديو للوسم. يغطي البرنامج دورة حياة البيانات الكاملة، بما في ذلك استيراد الأصول من التخزين المحلي أو السحابي وتحويل البيانات المعلقة إلى تنسيقات تعلم آلة متنوعة عبر تصديرات قائمة على المخططات (schema-based). كما يتضمن تشفيراً قائماً على الرموز لتأمين إعدادات تكوين المشروع الحساسة.

    Enables tagging of objects across video sequences by navigating through extracted frames to maintain temporal consistency.

    TypeScript
    عرض على GitHub↗4,427
السابق12التالي
  1. Home
  2. Graphics & Multimedia
  3. Media Processing and Analysis
  4. Media Manipulation
  5. Video Processing Tools

استكشف الوسوم الفرعية

  • Transcoding OffloadersMechanisms for delegating resource-intensive video encoding tasks to external hardware or software processes. **Distinct from Video Processing Tools:** Distinct from general video processing tools: focuses on the architectural delegation of encoding tasks to sidecar processes.
  • Video Frame AnnotationsThe process of adding labels and metadata to individual frames within a video sequence. **Distinct from Video Frame Navigators:** Distinct from Video Frame Navigators: focuses on the act of annotating the content of the frame, not the navigation tool.
  • Video Frame Navigators4 وسوم فرعيةTools that allow users to step through, seek, and inspect individual frames within a video file.