VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized benchmarks and measure model accuracy across diverse visual datasets. It serves as a multimodal model benchmark and performance toolkit for calculating metrics and comparing model responses. The toolkit includes a specialized visual reasoning evaluator that uses adversarial samples to distinguish actual image understanding from reliance on language patterns. It also provides capabilities for image generation evaluation, testing a model's ability to create or modify visuals based on tex
✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
🌐 Homepage | 🤗 Paper | 📖 arXiv | 🤗 Dataset | 🏆 EvalAI | GitHub
An automatic evaluation framework for Multimodal Chain-of-Thought.
paper: https://arxiv.org/abs/2307.02469 page: https://lynx-llm.github.io/
The main features of bytedance/lynx-llm are: Multimodal Benchmarks.
Open-source alternatives to bytedance/lynx-llm include: open-compass/vlmevalkit — VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized… bradyfu/video-mme — ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. cmmmu-benchmark/cmmmu — 🌐 Homepage | 🤗 Paper | 📖 arXiv | 🤗 Dataset | 🏆 EvalAI | GitHub. damo-nlp-sg/m3exam — Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models". dcdmllm/cheetah — Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions. alenai97/miceval — An automatic evaluation framework for Multimodal Chain-of-Thought.