awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
jayleicn avatar

jayleicn/ClipBERT

0
View on GitHub↗
730 stars·87 forks·Python·MIT·6 viewsarxiv.org/abs/2102.06183↗

ClipBERT

Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling

Features

  • Multimodal Pretraining - Sparse sampling for video-and-language representation learning.
  • Video Retrieval Models - Sparse sampling framework for video-and-language learning.
  • Video Understanding - Video-and-language learning via sparse sampling.

Star history

Star history chart for jayleicn/clipbertStar history chart for jayleicn/clipbert

How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does jayleicn/clipbert do?

Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling

What are the main features of jayleicn/clipbert?

The main features of jayleicn/clipbert are: Multimodal Pretraining, Video Retrieval Models, Video Understanding.

What are some open-source alternatives to jayleicn/clipbert?

Open-source alternatives to jayleicn/clipbert include: arrowluo/clip4clip — An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval". cryhanfang/clip2video — The implementation of paper CLIP2Video: Mastering Video-Text Retrieval via Image CLIP. google-research/big_vision — This project is a research framework and toolkit designed for training large-scale vision transformers and multimodal… evolvinglmms-lab/otter — Otter is a framework and toolkit for the pretraining, fine-tuning, and evaluation of vision-language models. It… facebookresearch/slowfast — SlowFast is a PyTorch video understanding framework and spatiotemporal neural network library. It serves as a toolset… llava-vl/llava-next — LLaVA-NeXT is a multimodal large language model framework and training toolkit designed to process interleaved images…

Open-source alternatives to ClipBERT

Similar open-source projects, ranked by how many features they share with ClipBERT.
  • cryhanfang/clip2videoCryhanFang avatar

    CryhanFang/CLIP2Video

    260View on GitHub↗

    The implementation of paper CLIP2Video: Mastering Video-Text Retrieval via Image CLIP.

    Python
    View on GitHub↗260
  • arrowluo/clip4clipArrowLuo avatar

    ArrowLuo/CLIP4Clip

    1,028View on GitHub↗

    An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"

    Pythonactivitynetclipdidemo
    View on GitHub↗1,028
  • evolvinglmms-lab/otterEvolvingLMMs-Lab avatar

    EvolvingLMMs-Lab/Otter

    3,331View on GitHub↗

    Otter is a framework and toolkit for the pretraining, fine-tuning, and evaluation of vision-language models. It provides a pipeline for training large language models to process high-resolution images and video frames, integrating visual encoders with textual token spaces. The system is designed for multi-visual input processing, allowing models to interpret multiple images or video sequences within a single prompt. It supports multi-round conversation management to maintain context across interactions for detailed scene comprehension and visual reasoning. The framework covers a full develop

    Pythonartificial-inteligencechatgptdeep-learning
    View on GitHub↗3,331
  • facebookresearch/slowfastfacebookresearch avatar

    facebookresearch/SlowFast

    7,377View on GitHub↗

    SlowFast is a PyTorch video understanding framework and spatiotemporal neural network library. It serves as a toolset for video action recognition, enabling the training and evaluation of models designed to classify complex activities and objects within video sequences. The framework is distinguished by its use of dual-pathway spatiotemporal sampling to capture both slow and fast motions. It supports self-supervised video learning for pre-training models on unlabeled data and employs multigrid spatiotemporal training to optimize learning across multiple spatial and temporal resolutions. The

    Python
    View on GitHub↗7,377
  • See all 30 alternatives to ClipBERT→