Our servers break again :(. I have updated the links so that they should work fine now. Sorry for the inconvenience. Please let me for any further issues. Thanks! --Hao, Dec 03
ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models
An Open-source Toolkit for LLM Development
This repo hosts the source code for our AAAI2020 work Vision-Language Pre-training (VLP). We have released the pre-trained model on Conceptual Captions dataset and fine-tuned models on COCO Captions and Flickr30k for image captioning and VQA 2.0 for VQA.
The main features of luoweizhou/vlp are: Vision Language Models.
Open-source alternatives to luoweizhou/vlp include: airsplay/lxmert — Our servers break again :(. I have updated the links so that they should work fine now. Sorry for the inconvenience.… alibaba/conv-llava — ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models. alpha-vllm/llama2-accessory — An Open-source Toolkit for LLM Development. apple/ml-aim — This repository provides the code and model checkpoints for AIMv1 and AIMv2 research projects. baaivision/eve — 2024/05: Unveiling Encoder-Free Vision-Language Models (NeurIPS 2024, spotlight). aidc-ai/parrot — 📍Quick Start • 👨🏫Acknowledgement • 🤗Contact.