How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.
VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized benchmarks and measure model accuracy across diverse visual datasets. It serves as a multimodal model benchmark and performance toolkit for calculating metrics and comparing model responses. The toolkit includes a specialized visual reasoning evaluator that uses adversarial samples to distinguish actual image understanding from reliance on language patterns. It also provides capabilities for image generation evaluation, testing a model's ability to create or modify visuals based on tex
β¨β¨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
paper: https://arxiv.org/abs/2307.02469 page: https://lynx-llm.github.io/
An automatic evaluation framework for Multimodal Chain-of-Thought.
π Homepage | π€ Paper | π arXiv | π€ Dataset | π EvalAI | GitHub
The main features of cmmmu-benchmark/cmmmu are: Multimodal Benchmarks.
Open-source alternatives to cmmmu-benchmark/cmmmu include: open-compass/vlmevalkit β VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardizedβ¦ bradyfu/video-mme β β¨β¨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. bytedance/lynx-llm β paper: https://arxiv.org/abs/2307.02469 page: https://lynx-llm.github.io/. damo-nlp-sg/m3exam β Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models". dcdmllm/cheetah β Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions. alenai97/miceval β An automatic evaluation framework for Multimodal Chain-of-Thought.