awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
FreedomIntelligence avatar

FreedomIntelligence/MLLM-Bench

0
View on GitHub↗
76 stars·4 forks·Python·13 views

MLLM Bench

MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria

Features

  • Evaluation Benchmarks - Evaluating multimodal models using automated scoring.
  • Multimodal Benchmarks - Evaluation framework using GPT-4V with per-sample criteria.

Star history

Star history chart for freedomintelligence/mllm-benchStar history chart for freedomintelligence/mllm-bench

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with MLLM Bench

These projects share indexed features with MLLM Bench. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • bradyfu/video-mmeBradyFU avatar

    BradyFU/Video-MME

    779View on GitHub↗

    ✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

    View on GitHub↗779
  • damo-nlp-sg/m3examDAMO-NLP-SG avatar

    DAMO-NLP-SG/M3Exam

    105View on GitHub↗

    Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models"

    Pythonai-educationchatgptevaluation
    View on GitHub↗105
  • ailab-cvc/seed-benchAILab-CVC avatar

    AILab-CVC/SEED-Bench

    364View on GitHub↗

    (CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions.

    Python
    View on GitHub↗364
  • gzcch/bingogzcch avatar

    gzcch/Bingo

    55View on GitHub↗

    Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, Huaxiu Yao *Equal Contribution

    View on GitHub↗55
Compare all 30 related projects→

Frequently asked questions

What does freedomintelligence/mllm-bench do?

MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria

What are the main features of freedomintelligence/mllm-bench?

The main features of freedomintelligence/mllm-bench are: Evaluation Benchmarks, Multimodal Benchmarks.

Which projects share features with freedomintelligence/mllm-bench?

Projects with overlapping indexed features include: damo-nlp-sg/m3exam — Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models". hypjudy/sparkles — Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models. bradyfu/video-mme — ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. ailab-cvc/seed-bench — (CVPR2024)A benchmark for evaluating Multimodal LLMs using multiple-choice questions. gzcch/bingo — Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, Huaxiu Yao *Equal Contribution. lerogo/mmgenbench — Official repository of MMGenBench.