18 open-source projects similar to actava-ai/chi-bench, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.
The code to visualize colon-bench (MICCAI 2026) data and run evaluations of MLLMs on the benchmark.
Go to https://mimic.physionet.org/ for access. Once you have the authority for the dataset, download the dataset at the data folder under the same directory as the this repository.
Abstract: Medical image analysis increasingly relies on large vision-language models (VLMs), yet most systems remain single-pass black boxes that offer limited control over reasoning, safety, and spatial grounding. We propose R^4, an agentic framework that decomposes medical imaging workflows…
The largest open-source medical AI skill library for OpenClaw.
MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning
ShortageSim is a comprehensive multi-agent simulation framework that models pharmaceutical supply chain dynamics during drug shortage events. By leveraging Large Language Models (LLMs) to power agent decision-making, the system captures realistic responses to regulatory signals and market…
Verifiable Epistemic Reasoning for Image-Derived Hypothesis Testing via Agentic Systems
MedBeads is an Immutable, Agent-Native Data Infrastructure designed to address the "Context Mismatch" in medical AI. By restructuring medical records from mutable relational databases into a Merkle Directed Acyclic Graph (DAG), MedBeads provides explicit causal linking, tamper-evidence, and…
A medical machine learning benchmark platform for evaluating automated machine learning agents on realistic healthcare tasks.
09/13/2024 🍓 We release new results and support for o1! - 08/17/2024 🎆 Major updates 🎇 - 🏥 A new suite of cases (AgentClinic-MIMIC-IV), based on real clinical cases from MIMIC-IV (requires approval from https://physionet.org/content/mimiciv/2.2/)! - More AgentClinic-MedQA cases 107 →…
DeepTumorVQA benchmark for VLMs and Agents (10k testing samples)
Observational studies can yield clinically actionable evidence at scale, but executing them on real-world databases is open-ended and requires coherent decisions across cohort construction, analysis, and reporting. Prior evaluations of LLM agents emphasize isolated steps or single answers,…
This repository contains implementation of MedAgentBench, and it is built on top of AgentBench. Please note that this code repo is intended for research purpose, and might not be suitable for large-scale production.
This benchmark system simulates an interactive conversation between a patient and an expert. The system evaluates how well participants' expert modules can handle realistic patient queries by either asking relevant questions or making final decisions based on the conversation history.
Access the paper and technical details here via arXiv IMAS is an advanced agentic medical assistant system designed to enhance healthcare delivery in rural areas, especially where experienced medical professionals are scarce. Leveraging fine-tuned healthcare domain-adapted Large Language Models…
ICLR'26 MedAgentGYM: Training LLM Agents for Code-Based Medical Reasoning at Scale
This repository provides a reference implementation of Med-Inquire and EvoClinician, including:
🎉 Our paper has been accepted to the NeurIPS 2025 Datasets & Benchmarks Track! 🎉