Code for the paper Benchmarking Large Multimodal Models against Common Corruptions.
The main features of sail-sg/mmcbench are: Evaluation Benchmarks, Multimodal Benchmarks.
Open-source alternatives to sail-sg/mmcbench include: damo-nlp-sg/m3exam — Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models". gzcch/bingo — Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, Huaxiu Yao *Equal Contribution. bradyfu/video-mme — ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. ailab-cvc/seed-bench — (CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions. freedomintelligence/mllm-bench — MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria. hypjudy/sparkles — Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models.
✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models"
(CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria