awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
MCG-NJU avatar

MCG-NJU/VideoMAE

0
View on GitHub↗
1,760 stars·168 forks·Python·17 viewsarxiv.org/abs/2203.12602↗

VideoMAE

[NeurIPS 2022 Spotlight] VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Features

  • Perception Models - Data-efficient self-supervised pre-training for video models.
  • Video Understanding and Tracking - Masked autoencoders for self-supervised video pre-training.

Star history

Star history chart for mcg-nju/videomaeStar history chart for mcg-nju/videomae

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with VideoMAE

These projects share indexed features with VideoMAE. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • anirudh257/strmAnirudh257 avatar

    Anirudh257/strm

    102View on GitHub↗

    CVPR 2022 Official Pytorch Implementation for "Spatio-temporal Relation Modeling for Few-shot Action Recognition". SOTA Results for Few-shot Action Recognition

    Python
    View on GitHub↗102
  • botaoye/ostrackbotaoye avatar

    botaoye/OSTrack

    653View on GitHub↗

    ECCV 2022 Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework

    Python
    View on GitHub↗653
  • chenfei-wu/taskmatrixchenfei-wu avatar

    chenfei-wu/TaskMatrix

    34,082View on GitHub↗

    TaskMatrix is a multimodal AI chat interface and visual task orchestrator. It combines language models with visual recognition to enable the exchange, analysis, and modification of images within a conversational environment. The system coordinates multiple foundation models through orchestration pipelines that chain language, detection, and segmentation models. This allows for complex visual operations, such as using text instructions to guide image masking and executing modular inpainting workflows to edit specific image regions. The project includes a computer vision toolset for object det

    Python
    View on GitHub↗34,082
  • aigc-audio/audiogptAIGC-Audio avatar

    AIGC-Audio/AudioGPT

    10,174View on GitHub↗

    AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural audio pipelines. It functions as a multimodal audio generator and processing system, integrating a collection of pretrained models to handle speech synthesis, sound generation, and audio manipulation. The system is distinguished by its ability to generate audio from diverse inputs, including text and images, and its capacity to produce synchronized talking head videos. It also operates as a neural speech translator, converting spoken language between different tongues while pre

    Pythonaudiogptmusic
    View on GitHub↗10,174
Compare all 25 related projects→

Frequently asked questions

What does mcg-nju/videomae do?

[NeurIPS 2022 Spotlight] VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

What are the main features of mcg-nju/videomae?

The main features of mcg-nju/videomae are: Perception Models, Video Understanding and Tracking.

Which projects share features with mcg-nju/videomae?

Projects with overlapping indexed features include: anirudh257/strm — [CVPR 2022] Official Pytorch Implementation for "Spatio-temporal Relation Modeling for Few-shot Action Recognition".… botaoye/ostrack — [ECCV 2022] Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework. chenfei-wu/taskmatrix — TaskMatrix is a multimodal AI chat interface and visual task orchestrator. It combines language models with visual… chenxin-dlut/transt — Transformer Tracking (CVPR2021). cvlab-columbia/viper — Code for the paper "ViperGPT: Visual Inference via Python Execution for Reasoning". aigc-audio/audiogpt — AudioGPT is an LLM-driven audio framework and processing suite that uses large language models to orchestrate neural…