This is a PyTorch self-supervised learning framework designed to train models that learn visual representations from video. It implements a joint-embedding predictive architecture that extracts spatio-temporal features by predicting missing regions of a signal within a latent representation space rather than reconstructing raw pixels. The project includes a latent space visualization tool that uses a conditional diffusion model to decode feature-space predictions back into pixels. This allows for the verification of learned representations by transforming abstract predictions into interpretab
This code implements the video- and sensor-based action segmentation models from Temporal Convolutional Networks for Action Segmentation and Detection by Colin Lea, Michael Flynn, Rene Vidal, Austin Reiter, Greg Hager arXiv 2016 (in-review).
NeurIPS 2021 Spotlight Official implementation of Long Short-Term Transformer for Online Action Detection
الميزات الرئيسية لـ bestjuly/iic هي: Video Representation Learning.
تشمل البدائل مفتوحة المصدر لـ bestjuly/iic: facebookresearch/jepa — This is a PyTorch self-supervised learning framework designed to train models that learn visual representations from… colincsl/temporalconvolutionalnetworks — This code implements the video- and sensor-based action segmentation models from Temporal Convolutional Networks for… decisionforce/vthcl. deepmind/kinetics-i3d — Convolutional neural network model for video classification trained on the Kinetics dataset. dmlc/decord — An efficient video loader for deep learning with smart shuffling that's super easy to digest. amazon-research/long-short-term-transformer — [NeurIPS 2021 Spotlight] Official implementation of Long Short-Term Transformer for Online Action Detection.