How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models"
(CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
Official repository of MMGenBench
The main features of lerogo/mmgenbench are: Evaluation Benchmarks, Multimodal Benchmarks.
Projects with overlapping indexed features include: damo-nlp-sg/m3exam — Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models". gzcch/bingo — Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, Huaxiu Yao *Equal Contribution. bradyfu/video-mme — ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. ailab-cvc/seed-bench — (CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions. freedomintelligence/mllm-bench — MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria. hypjudy/sparkles — Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models.