18 open-source projects similar to somewordstoolate/rwe-bench, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.
C linical H ealthcare I n-Situ Environment Benchmark for long-horizon, policy-rich healthcare workflow agents
The code to visualize colon-bench (MICCAI 2026) data and run evaluations of MLLMs on the benchmark.
Go to https://mimic.physionet.org/ for access. Once you have the authority for the dataset, download the dataset at the data folder under the same directory as the this repository.
Automated Drug Combination Extraction (DCE) from large-scale biomedical literature is important for precision medicine and pharmacological research. However, existing extraction methodologies predominantly focus on binary interactions. When tasked with complex, variable-length n-ary…
MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning
Aegle is a graph-based, multi-agent framework that virtualizes multidisciplinary clinical reasoning for outpatient intake. It coordinates an orchestrator, specialist agents, an aggregator, and a standardized patient to conduct evidence-grounded dialogue and generate structured Initial Progress…
Infherno is an end-to-end agent that transforms unstructured clinical notes into structured FHIR (Fast Healthcare Interoperability Resources) format. It automates the parsing and mapping of free-text medical documentation into standardized FHIR resources, enabling interoperability across…
2026-06-03 We add demo inference entry. Try your custom question with bash rundemoinference.sh. - 2026-03-26 We release code and paper: Unified-MAS: Universally Generating Domain-Specific Nodes for Empowering Automatic Multi-Agent Systems.
ICML 2026 Official codebase for From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAG.
EHRFlow is a large language model-driven platform designed to simplify electronic health record (EHR) data analysis for physicians through natural language interactions, eliminating the need for complex coding.
A medical machine learning benchmark platform for evaluating automated machine learning agents on realistic healthcare tasks.
09/13/2024 🍓 We release new results and support for o1! - 08/17/2024 🎆 Major updates 🎇 - 🏥 A new suite of cases (AgentClinic-MIMIC-IV), based on real clinical cases from MIMIC-IV (requires approval from https://physionet.org/content/mimiciv/2.2/)! - More AgentClinic-MedQA cases 107 →…
DeepTumorVQA benchmark for VLMs and Agents (10k testing samples)
Meissa is a multi-modal medical agent, built on trajectory-based agentic behavior distillation framework.
This repository contains implementation of MedAgentBench, and it is built on top of AgentBench. Please note that this code repo is intended for research purpose, and might not be suitable for large-scale production.
This benchmark system simulates an interactive conversation between a patient and an expert. The system evaluates how well participants' expert modules can handle realistic patient queries by either asking relevant questions or making final decisions based on the conversation history.
🎉 Our paper has been accepted to the NeurIPS 2025 Datasets & Benchmarks Track! 🎉