[NeurIPS 2024] MATH-Vision dataset and code to measure multimodal mathematical reasoning capabilities.
The main features of mathvision-cuhk/mathvision are: Multimodal Benchmarks.
Open-source alternatives to mathvision-cuhk/mathvision include: open-compass/vlmevalkit — VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized… bradyfu/video-mme — ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. bytedance/lynx-llm — paper: https://arxiv.org/abs/2307.02469 page: https://lynx-llm.github.io/. cmmmu-benchmark/cmmmu — 🌐 Homepage | 🤗 Paper | 📖 arXiv | 🤗 Dataset | 🏆 EvalAI | GitHub. damo-nlp-sg/m3exam — Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models". alenai97/miceval — An automatic evaluation framework for Multimodal Chain-of-Thought.
VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized benchmarks and measure model accuracy across diverse visual datasets. It serves as a multimodal model benchmark and performance toolkit for calculating metrics and comparing model responses. The toolkit includes a specialized visual reasoning evaluator that uses adversarial samples to distinguish actual image understanding from reliance on language patterns. It also provides capabilities for image generation evaluation, testing a model's ability to create or modify visuals based on tex
✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
paper: https://arxiv.org/abs/2307.02469 page: https://lynx-llm.github.io/
An automatic evaluation framework for Multimodal Chain-of-Thought.