16 dépôts
Computing the per-pixel movement between frames to support temporal rendering effects.
Distinct from Motion Blur Simulations: Different from general motion blur; it is the underlying calculation of vectors for TAA and blur.
Explore 16 awesome GitHub repositories matching graphics & multimedia · Motion Vector Calculation. Refine with filters or upvote what's useful.
This project is an educational suite and technical guide designed for mastering video codecs and signal processing. It provides a structured curriculum through an engineering course, interactive labs, and tutorials focused on the fundamental principles of video compression and digital signal processing. The resource includes a technical guide for analyzing specific codecs like AV1, VP9, and H.265. It distinguishes itself by providing a containerized media lab, which ensures a consistent development environment for experimenting with video technology tools and notebooks. The project covers a
Identifies and encodes changes between sequential video frames using block motion compensation and motion vectors.
This project is an open-source 3D game engine designed for building high-fidelity games, simulations, and cinematic environments. It functions as a robotics simulation platform with native integration for ROS 2 to model robot controllers and sensors. The engine features a multi-threaded Forward+ physically based renderer that supports hardware-accelerated ray tracing and global illumination. The system is built on a modular extension architecture using Gems to add or replace features without modifying core binaries. It includes a native SDK for AWS cloud integration, enabling IAM authenticati
Computes frame position differences in shaders to enable motion blur and temporal anti-aliasing.
Pose-animator is a system that maps real-time body and face tracking data to 2D vector illustrations. It functions as a skeletal animation engine and motion controller that translates human keypoint recognition into instantaneous SVG path updates. The project enables real-time motion capture from webcam feeds and pose extraction from static images. It utilizes a skeletal rig to link virtual bones to vector character surfaces, allowing for the animation of custom characters and interactive avatars. The tool incorporates client-side machine learning inference for processing camera frames, coor
Smooths the transitions between disparate ML detection results to prevent jittering in character motion.
jetson-inference is a set of libraries and tools for executing optimized deep learning models on embedded GPU hardware. Its primary purpose is to enable real-time computer vision and AI inference at the edge with low latency and high throughput. The project distinguishes itself through high-performance streaming analytics and the ability to execute concurrent AI pipelines on auto-grade silicon. It provides specialized support for multi-sensor stream processing, utilizing zero-copy data transport to load camera frames directly into GPU memory. The codebase covers a broad surface of capabiliti
Calculates relative pixel motion between frames using dedicated GPU hardware to track object movement.
SAMURAI is a zero-shot visual tracking model that adapts the Segment Anything architecture for video object segmentation. It uses a first-frame prompt, such as a bounding box or mask, to initialize tracking, then employs a motion-aware memory mechanism that stores and updates temporal motion features across frames to guide mask refinement. An online memory update strategy continuously refreshes this memory with new frame predictions, while temporal motion encoding computes optical flow between consecutive frames to inform object boundary and occlusion handling. The system is designed for real
Computes optical flow between consecutive frames to inform object boundary and occlusion handling.
StoryDiffusion is a generative AI system designed for consistent character image and video generation. It utilizes a pluggable cross-attention module to inject shared character representations into pretrained diffusion models, allowing for visual identity stability across multiple images and scenes without retraining the base model. The project features a video generation pipeline that produces temporally coherent sequences from text prompts or condition images. It employs a latent space motion interpolator to predict intermediate frames and semantic motion, enabling long-range video generati
Generates intermediate video frames by interpolating semantic data within a variational autoencoder.
Uses motion estimation to generate intermediate frames for smooth video playback.
OpenCVSharp is a .NET library that wraps native OpenCV functions, providing C# developers with access to OpenCV's computer vision capabilities through an API that mirrors the native C/C++ style. It serves as a managed wrapper for image processing, feature detection, object detection, and image manipulation tasks, while also handling automatic disposal of unmanaged OpenCV resources like Mat objects to prevent memory leaks in .NET applications. The library enables keypoint detection and descriptor extraction using algorithms such as AKAZE, BRISK, or FAST, with brute-force or FLANN-based matchin
Implements dense optical flow estimation using the Brox variational method for motion analysis.
ToonCrafter is a model that combines latent diffusion, reference-based colorization, and sketch-guided control for cartoon animation and interpolation. It functions as a cartoon video interpolation model, a reference-based colorization model, and a sketch-guided animation tool, all built on a latent diffusion animation framework. The project distinguishes itself by integrating three core capabilities into a single pipeline: generating smooth intermediate frames between two cartoon images using diffusion-based priors, transferring color and style from a reference image onto black-and-white ske
Generates intermediate cartoon frames by interpolating data within a compressed latent space.
AniPortrait est un pipeline de synthèse vidéo IA conçu pour générer des portraits parlants photoréalistes et des animations faciales. Il fonctionne comme un générateur de têtes parlantes et un animateur piloté par l'audio qui synchronise les mouvements des lèvres, les expressions et les poses de la tête avec la parole ou des sources vidéo de référence. Le système inclut un outil de transfert d'expression faciale pour réenacter les mouvements d'une vidéo source sur une image de référence statique. Il utilise un modèle de diffusion latente avec un conditionnement d'image basé sur la référence pour maintenir l'identité visuelle et la cohérence à travers les images générées. Le pipeline couvre le mappage audio-vers-expression, le contrôle de mouvement guidé par la pose et la synthèse vidéo photoréaliste. Il incorpore un suréchantillonnage par interpolation d'images pour accélérer le processus de génération et réduire le temps de rendu total.
Uses latent space interpolation to generate intermediate frames for smoother video and faster rendering.
MathUtilities est une collection de boîtes à outils spécialisées fournissant des moteurs pour la géométrie, la vision par ordinateur, les mathématiques, la simulation physique et le traitement du signal. Elle fonctionne comme une bibliothèque complète de mathématiques et de physique axée sur l'algèbre linéaire, l'optimisation numérique et les calculs géométriques pour les applications techniques. Le projet se distingue par une boîte à outils de simulation physique et un moteur de géométrie 3D. Ceux-ci fournissent des capacités pour l'intégration de Verlet, des solveurs de cinématique inverse itératifs, le rendu de champ de distance via raymarching volumétrique et la déformation de géométrie de maillage. Il inclut également un utilitaire de vision par ordinateur pour estimer le mouvement relatif de la caméra et générer des projections fisheye. La bibliothèque couvre de larges domaines de capacités, notamment les systèmes de détection de collision utilisant les différences de Minkowski et le hachage spatial, la planification de mouvement robotique et le traitement du signal pour la réduction du bruit utilisant des filtres de Kalman. Les fonctionnalités supplémentaires incluent l'optimisation de données numériques, les opérations d'algèbre linéaire pour l'ajustement de nuages de points et la sérialisation JSON basée sur la réflexion pour les hiérarchies d'objets.
Solves for the relative motion between two images by tracking the movement of feature points.
Auto-editor est un éditeur vidéo automatisé en ligne de commande qui utilise FFmpeg pour supprimer le silence et les séquences inactives des fichiers vidéo. Il fonctionne comme une suite de traitement avec des générateurs de coupes spécialisés qui identifient les segments à découper en fonction des seuils de volume, de l'analyse de mouvement et de la transcription parole-vers-texte. L'outil se distingue en offrant un flux de travail de post-production flexible, permettant aux utilisateurs d'exporter des timelines de coupes automatisées sous forme de fichiers XML ou JSON pour une utilisation dans des logiciels de montage non linéaire professionnels. Au-delà de la simple suppression, il peut effectuer des ajustements de lecture dynamiques, comme augmenter la vitesse des segments silencieux au lieu de les supprimer entièrement. Le projet couvre un large éventail de capacités de manipulation multimédia, incluant la normalisation audio, la réduction de la sibilance et des effets visuels comme la composition de calques, les superpositions graphiques et les transformations d'échelle. Il supporte également l'ingestion de médias distants via des URL et fournit des utilitaires pour prévisualiser les statistiques de montage sans rendre la vidéo finale.
Detects still footage by measuring pixel-level changes between frames to identify sections with minimal movement.
lsfg-vk is a Vulkan-based frame generation tool and graphics middleware designed to increase perceived frame rates in graphics applications. It implements a lossless scaling algorithm to insert generated frames between rendered ones, increasing motion smoothness without reducing image quality. The project features a dedicated programming interface for integrating frame generation logic into external applications and includes an application profile manager. This manager allows specific generation settings to be assigned to individual executables, automating configuration upon application start
Uses Vulkan-based motion vector analysis and pixel data to generate intermediate frames for smoother motion.
QualityScaler is an AI video upscaler and local media processing tool designed to increase the resolution and visual quality of videos and images. It uses deep learning models to enhance detail and remove noise, operating as an offline application that executes all computations on local hardware. The project functions as a GPU-accelerated media processor that distributes workloads across multiple graphics cards to increase rendering speed. To prevent memory overflow during high-resolution tasks, it employs a tiled image processing method that splits large assets into smaller sections. The sy
Implements motion-based frame interpolation to blend original and upscaled frames for smoother transitions.
Real-Video-Enhancer is a cross-platform desktop application that utilizes neural networks to upscale resolution, generate intermediate frames, and denoise video files. It functions as a deep learning video processor that runs restoration models through hardware acceleration, dispatching heavy prediction workloads directly to underlying graphics hardware. The software executes optical-flow-based frame interpolation to increase framerates and motion smoothness, alongside dedicated filtering models that remove digital noise and blocky compression artifacts from compressed video streams. Additio
Generates intermediate video frames using motion estimation and neural interpolation models to increase framerate and smoothness.
Bar chart race is a Python data visualization library that transforms ordered tabular time-series data into animated bar and line chart races. It operates as an extension for rendering dynamic charts that illustrate how rankings and values change over time. The library interpolates wide-format chronological tables into densely sampled frame sequences, calculating intermediate numeric values to produce fluid motion animations. It orchestrates iterative canvas redraws through a plotting backend while supporting external multimedia encoders to export compressed standard video files. Generated a
Calculates intermediate numeric bar lengths between discrete time periods using linear interpolation to produce fluid motion.