How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
mPLUG-Owl: The Powerful Multi-modal Large Language Model Family
The main features of x-plug/mplug-owl are: Multimodal Benchmarks, Multimodal Foundation Models.
Projects with overlapping indexed features include: open-compass/vlmevalkit — VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized… baaivision/emu3.5 — Native Multimodal Models are World Learners. bradyfu/video-mme — ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. bytedance/lynx-llm — paper: https://arxiv.org/abs/2307.02469 page: https://lynx-llm.github.io/. cmmmu-benchmark/cmmmu — 🌐 Homepage | 🤗 Paper | 📖 arXiv | 🤗 Dataset | 🏆 EvalAI | GitHub. alenai97/miceval — An automatic evaluation framework for Multimodal Chain-of-Thought.
VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized benchmarks and measure model accuracy across diverse visual datasets. It serves as a multimodal model benchmark and performance toolkit for calculating metrics and comparing model responses. The toolkit includes a specialized visual reasoning evaluator that uses adversarial samples to distinguish actual image understanding from reliance on language patterns. It also provides capabilities for image generation evaluation, testing a model's ability to create or modify visuals based on tex
Native Multimodal Models are World Learners
✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
An automatic evaluation framework for Multimodal Chain-of-Thought.