How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models"
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, Huaxiu Yao *Equal Contribution
(CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.
The main features of ailab-cvc/seed-bench are: Evaluation Benchmarks, Multimodal Benchmarks.
Projects with overlapping indexed features include: hypjudy/sparkles — Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models. gzcch/bingo — Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, Huaxiu Yao *Equal Contribution. bradyfu/video-mme — ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. damo-nlp-sg/m3exam — Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models". freedomintelligence/mllm-bench — MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria. lerogo/mmgenbench — Official repository of MMGenBench.