awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to freedomintelligence/mllm-bench

Open-source alternatives to MLLM Bench

30 open-source projects similar to freedomintelligence/mllm-bench, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best MLLM Bench alternative.

  • ailab-cvc/seed-benchAILab-CVC avatar

    AILab-CVC/SEED-Bench

    364View on GitHub↗

    (CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.

    Python
    View on GitHub↗364
  • bradyfu/video-mmeBradyFU avatar

    BradyFU/Video-MME

    779View on GitHub↗

    ✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

    View on GitHub↗779
  • gzcch/bingogzcch avatar

    gzcch/Bingo

    55View on GitHub↗

    Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, Huaxiu Yao *Equal Contribution

    View on GitHub↗55
  • hypjudy/sparklesHYPJUDY avatar

    HYPJUDY/Sparkles

    45View on GitHub↗

    Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models

    Python
    View on GitHub↗45
  • tsb0601/mmvptsb0601 avatar

    tsb0601/MMVP

    363View on GitHub↗

    Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, Saining Xie

    Python
    View on GitHub↗363
  • open-compass/mmbenchopen-compass avatar

    open-compass/MMBench

    303View on GitHub↗

    Official Repo of "MMBench: Is Your Multi-modal Model an All-around Player?"

    View on GitHub↗303

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • lerogo/mmgenbenchlerogo avatar

    lerogo/MMGenBench

    119View on GitHub↗

    Official repository of MMGenBench

    Pythonllms-benchmarkingmllmmmgenbench
    View on GitHub↗119
  • opengvlab/multi-modality-arenaOpenGVLab avatar

    OpenGVLab/Multi-Modality-Arena

    558View on GitHub↗
    Pythonchatchatbotchatgpt
    View on GitHub↗558
  • openm3d/m3dbenchOpenM3D avatar

    OpenM3D/M3DBench

    61View on GitHub↗

    ECCV 2024 M3DBench introduces a comprehensive 3D instruction-following dataset with support for interleaved multi-modal prompts.

    Python3ddatasetinstruction-tuning
    View on GitHub↗61
  • yuweihao/mm-vetyuweihao avatar

    yuweihao/MM-Vet

    327View on GitHub↗

    MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities (ICML 2024)

    Python
    View on GitHub↗327
  • damo-nlp-sg/m3examDAMO-NLP-SG avatar

    DAMO-NLP-SG/M3Exam

    105View on GitHub↗

    Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models"

    Pythonai-educationchatgptevaluation
    View on GitHub↗105
  • sail-sg/mmcbenchsail-sg avatar

    sail-sg/MMCBench

    27View on GitHub↗

    Code for the paper Benchmarking Large Multimodal Models against Common Corruptions.

    Python
    View on GitHub↗27
  • openlamm/lammOpenLAMM avatar

    OpenLAMM/LAMM

    317View on GitHub↗

    NeurIPS 2023 Datasets and Benchmarks Track LAMM: Multi-Modal Large Language Models and Applications as AI Agents

    Python
    View on GitHub↗317
  • open-compass/vlmevalkitopen-compass avatar

    open-compass/VLMEvalKit

    3,824View on GitHub↗

    VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized benchmarks and measure model accuracy across diverse visual datasets. It serves as a multimodal model benchmark and performance toolkit for calculating metrics and comparing model responses. The toolkit includes a specialized visual reasoning evaluator that uses adversarial samples to distinguish actual image understanding from reliance on language patterns. It also provides capabilities for image generation evaluation, testing a model's ability to create or modify visuals based on tex

    Pythonchatgptclaudeclip
    View on GitHub↗3,824
  • pyspur-dev/pyspurPySpur-Dev avatar

    PySpur-Dev/pyspur

    5,677View on GitHub↗
    TypeScriptagentagentsai
    View on GitHub↗5,677
  • datawhalechina/prompt-engineering-for-developersdatawhalechina avatar

    datawhalechina/prompt-engineering-for-developers

    24,267View on GitHub↗

    This project is a technical curriculum and development guide focused on large language model prompt engineering, fine-tuning, and the creation of retrieval augmented generation applications. It serves as a comprehensive resource for developers to master crafting precise instructions and textual patterns to improve the quality and predictability of model outputs. The material covers the end-to-end workflow of adapting open-source models to specific datasets and integrating language models with vector databases to generate responses based on private information. It also provides a systematic ap

    Jupyter Notebook
    View on GitHub↗24,267
  • atsumiyai/updAtsuMiyai avatar

    AtsuMiyai/UPD

    82View on GitHub↗

    🤗 Dataset | 📖 arXiv | GitHub Atsuyuki Miyai 1   Jingkang Yang 2   Jingyang Zhang 3   Yifei Ming 4   Qing Yu 1,5   Go Irie 6   Sharon Yixuan Li 4   Hai Li 3   Ziwei Liu 2 Kiyoharu Aizawa 1 1 The University of Tokyo  2 S-Lab, Nanyang Technological…

    Python
    View on GitHub↗82
  • atr-dbi/scanqaATR-DBI avatar

    ATR-DBI/ScanQA

    160View on GitHub↗

    This is the official repository of our paper ScanQA: 3D Question Answering for Spatial Scene Understanding (CVPR 2022) by Daichi Azuma, Taiki Miyanishi, Shuhei Kurita, and Motoki Kawanabe. We propose a new 3D spatial understanding task for 3D question answering (3D-QA). In the 3D-QA task, models…

    Python
    View on GitHub↗160
  • alenai97/micevalalenai97 avatar

    alenai97/MiCEval

    6View on GitHub↗

    An automatic evaluation framework for Multimodal Chain-of-Thought.

    Python
    View on GitHub↗6
  • aoidragon/popeAoiDragon avatar

    AoiDragon/POPE

    118View on GitHub↗

    This repo provides the source code & data of our paper: Evaluating Object Hallucination in Large Vision-Language Models (EMNLP 2023).

    Python
    View on GitHub↗118
  • controllability/jailbreak-evaluationcontrollability avatar

    controllability/jailbreak-evaluation

    27View on GitHub↗

    The jailbreak-evaluation is an easy-to-use Python package for language model jailbreak evaluation. The jailbreak-evaluation is designed for comprehensive and accurate evaluation of language model jailbreak attempts. Currently, jailbreak-evaluation support evaluating a language model jailbreak…

    Python
    View on GitHub↗27
  • andreamaduzzi/crosscoherenceAndreAmaduzzi avatar

    AndreAmaduzzi/CrossCoherence

    1View on GitHub↗

    Train-val-test splits of GPT2Shape dataset can be found in folder gpt2shape: train val test

    View on GitHub↗1
  • albertwy/gpt-4v-evaluationalbertwy avatar

    albertwy/GPT-4V-Evaluation

    11View on GitHub↗

    Data for evaluating GPT-4V

    View on GitHub↗11
  • damo-nlp-sg/cmmDAMO-NLP-SG avatar

    DAMO-NLP-SG/CMM

    54View on GitHub↗

    🍎 Project Page 📖 arXiv Paper 📊 Dataset🏆 Leaderboard

    Python
    View on GitHub↗54
  • ai45lab/openrtAI45Lab avatar

    AI45Lab/OpenRT

    257View on GitHub↗

    Open-source red teaming framework for MLLMs with 42+ attack methods

    Python
    View on GitHub↗257
  • cmmmu-benchmark/cmmmuCMMMU-Benchmark avatar

    CMMMU-Benchmark/CMMMU

    48View on GitHub↗

    🌐 Homepage | 🤗 Paper | 📖 arXiv | 🤗 Dataset | 🏆 EvalAI | GitHub

    Python
    View on GitHub↗48
  • daveredrum/scan2capdaveredrum avatar

    daveredrum/Scan2Cap

    106View on GitHub↗

    We introduce the task of dense captioning in 3D scans from commodity RGB-D sensors. As input, we assume a point cloud of a 3D scene; the expected output is the bounding boxes along with the descriptions for the underlying objects. To address the 3D object detection and description problems, we…

    Python
    View on GitHub↗106
  • dcdmllm/cheetahDCDmllm avatar

    DCDmllm/Cheetah

    354View on GitHub↗

    Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

    Python
    View on GitHub↗354
  • declare-lab/red-instructdeclare-lab avatar

    declare-lab/red-instruct

    111View on GitHub↗

    Paper | Github | Dataset | Model

    Python
    View on GitHub↗111
  • chenllliang/mmevalprochenllliang avatar

    chenllliang/MMEvalPro

    25View on GitHub↗

    NAACL 2025 Source code for MMEvalPro, a more trustworthy and efficient benchmark for evaluating LMMs

    Python
    View on GitHub↗25