How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
This is the official PyTorch implementation of the ALBEF paper [Blog] . This repository supports pre-training on custom datasets, as well as finetuning on VQA, SNLI-VE, NLVR2, Image-Text Retrieval on MSCOCO and Flickr30k, and visual grounding on RefCOCO+. Pre-trained and finetuned checkpoints…
The main features of salesforce/albef are: Multimodal Pretraining, Vision Language Models.
Open-source alternatives to salesforce/albef include: jackroos/vl-bert — Code for ICLR 2020 paper "VL-BERT: Pre-training of Generic Visual-Linguistic Representations". uclanlp/visualbert. airsplay/lxmert — Our servers break again :(. I have updated the links so that they should work fine now. Sorry for the inconvenience.… evolvinglmms-lab/otter — Otter is a framework and toolkit for the pretraining, fine-tuning, and evaluation of vision-language models. It… google-research/big_vision — This project is a research framework and toolkit designed for training large-scale vision transformers and multimodal… apple/ml-aim — This repository provides the code and model checkpoints for AIMv1 and AIMv2 research projects.
Code for ICLR 2020 paper "VL-BERT: Pre-training of Generic Visual-Linguistic Representations".
Our servers break again :(. I have updated the links so that they should work fine now. Sorry for the inconvenience. Please let me for any further issues. Thanks! --Hao, Dec 03
This project is a research framework and toolkit designed for training large-scale vision transformers and multimodal language models. It provides a comprehensive suite for vision-language pretraining, enabling the development of models that map images and text into shared latent spaces. The framework is distinguished by its capabilities in high-fidelity image generation and multimodal research, utilizing normalizing flows and variational autoencoders to produce images from text prompts or class labels. It supports the development of both generative and contrastive models, allowing for a wide