Contextual Object Detection with Multimodal Large Language Models
The main features of yuhangzang/contextdet are: Model Utilities.
Open-source alternatives to yuhangzang/contextdet include: ailab-cvc/seed — Official implementation of SEED-LLaMA (ICLR 2024). dvlab-research/prompt-highlighter — [CVPR 2024] Prompt Highlighter: Interactive Control for Multi-Modal LLMs. huawei-noah/efficient-computing — Efficient computing methods developed by Huawei Noah's Ark Lab. kohjingyu/fromage — 🧀 Code and models for the ICML 2023 paper "Grounding Language Models to Images for Multimodal Inputs and Outputs". kohjingyu/gill — 🐟 Code and models for the NeurIPS 2023 paper "Generating Images with Multimodal Language Models". shi-labs/vcoder — [CVPR 2024] VCoder: Versatile Vision Encoders for Multimodal Large Language Models.
CVPR 2024 Prompt Highlighter: Interactive Control for Multi-Modal LLMs
Efficient computing methods developed by Huawei Noah's Ark Lab
🧀 Code and models for the ICML 2023 paper "Grounding Language Models to Images for Multimodal Inputs and Outputs".