30 open-source projects similar to cmmmu-benchmark/cmmmu, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.
VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized benchmarks and measure model accuracy across diverse visual datasets. It serves as a multimodal model benchmark and performance toolkit for calculating metrics and comparing model responses. The toolkit includes a specialized visual reasoning evaluator that uses adversarial samples to distinguish actual image understanding from reliance on language patterns. It also provides capabilities for image generation evaluation, testing a model's ability to create or modify visuals based on tex
An automatic evaluation framework for Multimodal Chain-of-Thought.
✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
paper: https://arxiv.org/abs/2307.02469 page: https://lynx-llm.github.io/
Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models"
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
ICLR'24 Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
NAACL 2024 MMC: Advancing Multimodal Chart Understanding with LLM Instruction Tuning
Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, Huaxiu Yao *Equal Contribution
Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models
NeurIPS 2025 The official repository of "Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning"
VELOCITI Benchmark Evaluation and Visualisation Code
Official repository of MMGenBench
@Author: Qiguang Chen @LastEditors: Qiguang Chen @Date: 2024-05-23 20:24:16 @LastEditTime: 2024-05-26 18:09:00 @Description: -->
ACL 2024 Findings "TempCompass: Do Video LLMs Really Understand Videos?", Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, Lu Hou
NeurIPS 2024 MATH-Vision dataset and code to measure multimodal mathematical reasoning capabilities.
Official implementation of "Visually Dehallucinative Instruction Generation: Know What You Don't Know"
Official Repo of "MMBench: Is Your Multi-modal Model an All-around Player?"
CVPR2024 HighlightVideoChatGPT ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.
NeurIPS 2023 Datasets and Benchmarks Track LAMM: Multi-Modal Large Language Models and Applications as AI Agents
ECCV 2024 M3DBench introduces a comprehensive 3D instruction-following dataset with support for interleaved multi-modal prompts.
A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models!
Code for the paper Benchmarking Large Multimodal Models against Common Corruptions.
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, Saining Xie
This is the repo for the paper SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval.
mPLUG-Owl: The Powerful Multi-modal Large Language Model Family
ICLR 2025 VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning
(CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.