How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.
The code to visualize colon-bench (MICCAI 2026) data and run evaluations of MLLMs on the benchmark.
MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning
A medical machine learning benchmark platform for evaluating automated machine learning agents on realistic healthcare tasks.
C linical H ealthcare I n-Situ Environment Benchmark for long-horizon, policy-rich healthcare workflow agents
Go to https://mimic.physionet.org/ for access. Once you have the authority for the dataset, download the dataset at the data folder under the same directory as the this repository.
The main features of clibench/clibench are: Healthcare Agent Benchmarks.
Projects with overlapping indexed features include: actava-ai/chi-bench — C linical H ealthcare I n-Situ Environment Benchmark for long-horizon, policy-rich healthcare workflow agents. ajhamdi/colon-bench-eval — The code to visualize colon-bench (MICCAI 2026) data and run evaluations of MLLMs on the benchmark. gersteinlab/medagents-benchmark — MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning. rajpurkarlab/rex-mle — A medical machine learning benchmark platform for evaluating automated machine learning agents on realistic healthcare… samuelschmidgall/agentclinic — [09/13/2024] 🍓 We release new results and support for o1! - [08/17/2024] 🎆 Major updates 🎇 - 🏥 A new suite of… schuture/deeptumorvqa — DeepTumorVQA benchmark for VLMs and Agents (10k testing samples).