awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to jayleicn/clipbert

Open-source alternatives to ClipBERT

30 open-source projects similar to jayleicn/clipbert, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best ClipBERT alternative.

  • arrowluo/clip4clipAvatar de ArrowLuo

    ArrowLuo/CLIP4Clip

    1,028Ver en GitHub↗

    An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"

    Pythonactivitynetclipdidemo
    Ver en GitHub↗1,028
  • cryhanfang/clip2videoAvatar de CryhanFang

    CryhanFang/CLIP2Video

    260Ver en GitHub↗

    The implementation of paper CLIP2Video: Mastering Video-Text Retrieval via Image CLIP.

    Python
    Ver en GitHub↗260
  • llava-vl/llava-nextAvatar de LLaVA-VL

    LLaVA-VL/LLaVA-NeXT

    4,695Ver en GitHub↗

    LLaVA-NeXT is a multimodal large language model framework and training toolkit designed to process interleaved images and video sequences to generate text. It functions as a visual language model that combines vision encoders with language models to perform complex reasoning, question answering, and video understanding. The system is capable of analyzing high-resolution images and temporal video frames to describe events, summarize actions, and reason across multiple visual inputs. It supports the interpretation of documents and charts, spatial environment analysis, and the generation of desc

    Python
    Ver en GitHub↗4,695
  • google-research/big_visionAvatar de google-research

    google-research/big_vision

    3,363Ver en GitHub↗

    This project is a research framework and toolkit designed for training large-scale vision transformers and multimodal language models. It provides a comprehensive suite for vision-language pretraining, enabling the development of models that map images and text into shared latent spaces. The framework is distinguished by its capabilities in high-fidelity image generation and multimodal research, utilizing normalizing flows and variational autoencoders to produce images from text prompts or class labels. It supports the development of both generative and contrastive models, allowing for a wide

    Jupyter Notebook
    Ver en GitHub↗3,363

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Find more with AI search
  • facebookresearch/slowfastAvatar de facebookresearch

    facebookresearch/SlowFast

    7,377Ver en GitHub↗

    SlowFast is a PyTorch video understanding framework and spatiotemporal neural network library. It serves as a toolset for video action recognition, enabling the training and evaluation of models designed to classify complex activities and objects within video sequences. The framework is distinguished by its use of dual-pathway spatiotemporal sampling to capture both slow and fast motions. It supports self-supervised video learning for pre-training models on unlabeled data and employs multigrid spatiotemporal training to optimize learning across multiple spatial and temporal resolutions. The

    Python
    Ver en GitHub↗7,377
  • evolvinglmms-lab/otterAvatar de EvolvingLMMs-Lab

    EvolvingLMMs-Lab/Otter

    3,331Ver en GitHub↗

    Otter is a framework and toolkit for the pretraining, fine-tuning, and evaluation of vision-language models. It provides a pipeline for training large language models to process high-resolution images and video frames, integrating visual encoders with textual token spaces. The system is designed for multi-visual input processing, allowing models to interpret multiple images or video sequences within a single prompt. It supports multi-round conversation management to maintain context across interactions for detailed scene comprehension and visual reasoning. The framework covers a full develop

    Pythonartificial-inteligencechatgptdeep-learning
    Ver en GitHub↗3,331
  • facebookresearch/vjepa2Avatar de facebookresearch

    facebookresearch/vjepa2

    3,021Ver en GitHub↗

    vjepa2 is a joint-embedding predictive architecture and video self-supervised learning framework. It functions as a visual representation learner and a robotic manipulation model designed to learn representations by predicting future latent states without reconstructing pixels. The system enables the pretraining of video encoders that learn temporally consistent features through masked-token prediction and multi-modal tokenization. It further maps these latent embeddings to specific physical movements via action-conditioned post-training to plan and execute robot arm grasping and picking task

    Python
    Ver en GitHub↗3,021
  • open-mmlab/mmaction2Avatar de open-mmlab

    open-mmlab/mmaction2

    5,066Ver en GitHub↗

    mmaction2 is a PyTorch video understanding toolbox designed for training and evaluating deep learning models. It serves as a framework for action recognition, temporal localization, and spatio-temporal action detection, providing specialized tools for both pixel-based video analysis and skeleton-based action recognition. The project distinguishes itself through a modular architecture featuring registry-based component discovery and hierarchical, config-driven model assembly. It supports multi-modal feature fusion, integrating RGB frames, optical flow, and audio, and includes capabilities for

    Python
    Ver en GitHub↗5,066
  • opengvlab/internvlAvatar de OpenGVLab

    OpenGVLab/InternVL

    10,061Ver en GitHub↗

    InternVL is a vision-language model framework that fuses a visual encoder with a large language model to translate image features into textual tokens for reasoning. It provides a system for multimodal inference and dialogue, enabling the processing of images and text to answer questions or generate descriptions. The project is distinguished by its high-resolution image processing, which uses dynamic tiling to maintain detail for images up to 4K resolution, and its chain-of-thought visual reasoning for solving complex mathematical and spatial problems. It also supports temporal frame sampling

    Pythongptgpt-4ogpt-4v
    Ver en GitHub↗10,061
  • gingsi/coot-videotextAvatar de gingsi

    gingsi/coot-videotext

    291Ver en GitHub↗

    2022-06-24 (v0.3.4): Bugfixes.

    Python
    Ver en GitHub↗291
  • huiguanlab/ms-slAvatar de HuiGuanLab

    HuiGuanLab/ms-sl

    57Ver en GitHub↗

    Source code of our ACM MM'2022 paper Partially Relevant Video Retrieval.

    Python
    Ver en GitHub↗57
  • huiguanlab/nrccrAvatar de HuiGuanLab

    HuiGuanLab/nrccr

    13Ver en GitHub↗

    source code of our paper Cross-Lingual Cross-Modal Retrieval with Noise-Robust Learning

    Python
    Ver en GitHub↗13
  • jackroos/vl-bertAvatar de jackroos

    jackroos/VL-BERT

    745Ver en GitHub↗

    Code for ICLR 2020 paper "VL-BERT: Pre-training of Generic Visual-Linguistic Representations".

    Jupyter Notebook
    Ver en GitHub↗745
  • jamespark3922/video-lang-contrast-setAvatar de jamespark3922

    jamespark3922/video-lang-contrast-set

    0Ver en GitHub↗

    Repository for Exposing the Limits of Video-Text Models through Contrast Sets (NAACL Short 2022).

    Shell
    Ver en GitHub↗0
  • kywen1119/cookieK

    kywen1119/COOKIE

    0Ver en GitHub↗

    Codes for ICCV2021 paper "COOKIE: Contrastive Cross-Modal Knowledge Sharing Pre-training for Vision-Language Representation". [PDF](https://openaccess.thecvf.com/content/ICCV2021/html/WenCOOKIEContrastiveCross-ModalKnowledgeSharingPre-TrainingforVision-LanguageRepresentationICCV2021paper.html)

    Ver en GitHub↗0
  • layer6ai-labs/xpoolAvatar de layer6ai-labs

    layer6ai-labs/xpool

    137Ver en GitHub↗

    X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval Satya Krishna Gorti , Noël Vouitsis , Junwei Ma* , Keyvan Golestan , Maksims Volkovs , Animesh Garg , Guangwei Yu

    Python
    Ver en GitHub↗137
  • leeyn-43/cloverAvatar de LeeYN-43

    LeeYN-43/Clover

    40Ver en GitHub↗

    Offical PyTorch implementation of Clover: Towards A Unified Video-Language Alignment and Fusion Model [arXiv](https://arxiv.org/abs/2207.07885)

    Python
    Ver en GitHub↗40
  • li-xirong/w2vvppAvatar de li-xirong

    li-xirong/w2vvpp

    29Ver en GitHub↗

    W2VV++: A fully deep learning solution for ad-hoc video search. The code assumes video-level CNN features have been extracted.

    Python
    Ver en GitHub↗29
  • lijiabei-7/rivrlAvatar de LiJiaBei-7

    LiJiaBei-7/rivrl

    19Ver en GitHub↗

    Source code of our paper Reading-strategy Inspired Visual Representation Learning for Text-to-Video Retrieval.

    Python
    Ver en GitHub↗19
  • m-bain/frozen-in-timeAvatar de m-bain

    m-bain/frozen-in-time

    376Ver en GitHub↗

    A Joint Video and Image Encoder for End-to-End Retrieval project page | paper | dataset | demo Repository containing the code, models, data for end-to-end retrieval. WebVid data can be found here

    Python
    Ver en GitHub↗376
  • microsoft/lavenderAvatar de microsoft

    microsoft/LAVENDER

    62Ver en GitHub↗

    Paper | Slide | Poster | Video

    Python
    Ver en GitHub↗62
  • mwray/semantic-video-retrievalM

    mwray/Semantic-Video-Retrieval

    0Ver en GitHub↗

    This repo contains code to evaluate for the semantic similarity video retrieval task, including: An example to generate a pandas dataframe from json annotations for YouCook2. A script to parse the captions using spacy. An optional script to create synset information using WordNet features. A…

    Ver en GitHub↗0
  • opengvlab/efficient-video-recognitionAvatar de OpenGVLab

    OpenGVLab/efficient-video-recognition

    184Ver en GitHub↗

    This is the official implementation of the paper Frozen CLIP models are Efficient Video Learners

    Python
    Ver en GitHub↗184
  • zhegan27/villaAvatar de zhegan27

    zhegan27/VILLA

    119Ver en GitHub↗

    This is the official repository of VILLA (NeurIPS 2020 Spotlight). This repository currently supports adversarial finetuning of UNITER on VQA, VCR, NLVR2, and SNLI-VE. Adversarial pre-training with in-domain data will be available soon. Both VILLA-base and VILLA-large pre-trained checkpoints are…

    Python
    Ver en GitHub↗119
  • airsplay/vokenizationAvatar de airsplay

    airsplay/vokenization

    191Ver en GitHub↗

    PyTorch code for EMNLP 2020 Paper "Vokenization: Improving Language Understanding with Visual Supervision"

    Python
    Ver en GitHub↗191
  • albanie/collaborative-expertsA

    albanie/collaborative-experts

    0Ver en GitHub↗

    This repo provides code: - TeachText which leverages complementary cues from multiple text encoders to provide an enhanced supervisory signal to the retrieval model using a generalize distillation setup (paper, project page) - Learning and evaluating joint video-text embeddings for the task of…

    Ver en GitHub↗0
  • antoine77340/howto100mAvatar de antoine77340

    antoine77340/howto100m

    303Ver en GitHub↗

    This repo provides code from the HowTo100M paper. We provide implementation of: - Our training procedure on HowTo100M for learning a joint text-video embedding - Our evaluation code on MSR-VTT, YouCook2 and LSMDC for Text-to-Video retrieval - A pretrain model on HowTo100M - Feature extraction…

    Python
    Ver en GitHub↗303
  • antoine77340/mixture-of-embedding-expertsAvatar de antoine77340

    antoine77340/Mixture-of-Embedding-Experts

    122Ver en GitHub↗

    This github repo provides a Pytorch implementation of the Mixture-of-Embeddings-Experts model (MEE) 1.

    Python
    Ver en GitHub↗122
  • bighuang624/vopAvatar de bighuang624

    bighuang624/VoP

    38Ver en GitHub↗

    News: The paper has been accepted to CVPR 2023!

    Ver en GitHub↗38
  • bryant1410/fitclipAvatar de bryant1410

    bryant1410/fitclip

    8Ver en GitHub↗

    This repo contains the code for the BMVC 2022 paper FitCLIP: Refining Large-Scale Pretrained Image-Text Models for Zero-Shot Video Understanding Tasks.

    Python
    Ver en GitHub↗8