awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
PKU-YuanGroup avatar

PKU-YuanGroup/Video-Bench

0
View on GitHub↗
140 stars·3 forks·Python·10 views

Video Bench

A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models!

Features

  • Multimodal Benchmarks - Benchmark and toolkit for video-based model evaluation.

Star history

Star history chart for pku-yuangroup/video-benchStar history chart for pku-yuangroup/video-bench

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Video Bench

These projects share indexed features with Video Bench. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • open-compass/vlmevalkitopen-compass avatar

    open-compass/VLMEvalKit

    3,824View on GitHub↗

    VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized benchmarks and measure model accuracy across diverse visual datasets. It serves as a multimodal model benchmark and performance toolkit for calculating metrics and comparing model responses. The toolkit includes a specialized visual reasoning evaluator that uses adversarial samples to distinguish actual image understanding from reliance on language patterns. It also provides capabilities for image generation evaluation, testing a model's ability to create or modify visuals based on tex

    Pythonchatgptclaudeclip
    View on GitHub↗3,824
  • bradyfu/video-mmeBradyFU avatar

    BradyFU/Video-MME

    779View on GitHub↗

    ✨✨CVPR 2025 Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

    View on GitHub↗779
  • bytedance/lynx-llmbytedance avatar

    bytedance/lynx-llm

    272View on GitHub↗

    paper: https://arxiv.org/abs/2307.02469 page: https://lynx-llm.github.io/

    Pythonresearch
    View on GitHub↗272
  • alenai97/micevalalenai97 avatar

    alenai97/MiCEval

    6View on GitHub↗

    An automatic evaluation framework for Multimodal Chain-of-Thought.

    Python
    View on GitHub↗6
Compare all 30 related projects→

Frequently asked questions

What does pku-yuangroup/video-bench do?

A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models!

What are the main features of pku-yuangroup/video-bench?

The main features of pku-yuangroup/video-bench are: Multimodal Benchmarks.

Which projects share features with pku-yuangroup/video-bench?

Projects with overlapping indexed features include: open-compass/vlmevalkit — VLMEvalKit is a vision-language model evaluation framework and inference engine designed to run standardized… bradyfu/video-mme — ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. bytedance/lynx-llm — paper: https://arxiv.org/abs/2307.02469 page: https://lynx-llm.github.io/. cmmmu-benchmark/cmmmu — 🌐 Homepage | 🤗 Paper | 📖 arXiv | 🤗 Dataset | 🏆 EvalAI | GitHub. damo-nlp-sg/m3exam — Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models". alenai97/miceval — An automatic evaluation framework for Multimodal Chain-of-Thought.