How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models"
(CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, Saining Xie
The main features of tsb0601/mmvp are: Evaluation Benchmarks, Multimodal Benchmarks.
Open-source alternatives to tsb0601/mmvp include: damo-nlp-sg/m3exam — Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models". gzcch/bingo — Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, Huaxiu Yao *Equal Contribution. bradyfu/video-mme — ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. ailab-cvc/seed-bench — (CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions. freedomintelligence/mllm-bench — MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria. hypjudy/sparkles — Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models.