28 open-source projects similar to gersteinlab/medagents-benchmark, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.
This repository contains implementation of MedAgentBench, and it is built on top of AgentBench. Please note that this code repo is intended for research purpose, and might not be suitable for large-scale production.
OpenEMR is an open-source electronic health record (EHR) system that also functions as a medical practice management platform and a patient portal, all integrated with standards-based health data exchange. It stores and manages patient health records, handles clinical workflows, supports scheduling and billing, and provides patients with secure self-service access to their information. Interoperability is built in via FHIR and C-CDA for exchanging records with external systems and Direct protocol for encrypted provider messaging. The system is designed to be extensible, with a modular plugin
Huatuo-Llama-Med-Chinese is a medical large language model specialized in processing and generating natural language text in Chinese. It is an instruction-tuned system designed to answer professional healthcare questions by leveraging a dedicated medical knowledge base. The model integrates structured medical literature and knowledge graphs to ensure clinical accuracy during response generation. It employs knowledge-graph augmented inference to combine structured entity relationships with neural network outputs. The system is developed through domain-specific weight adaptation, cross-lingual
This project is a collection of specialized toolkits and an agent skill library designed to equip large language model agents with the capabilities to perform complex scientific research across biology, chemistry, medicine, and physics. It provides a structured framework of integration paths and tools that allow agents to execute multi-step research workflows. The system is distinguished by its domain-specific toolsets, including a bioinformatics toolkit for genomic and single-cell analysis, a cheminformatics toolset for drug-target binding and lead compound optimization, and a multi-omics an
This is the official repository for "REFLECTOOL: Towards Reflection-Aware Tool-Augmented Clinical Agents"
MedRAX: Medical Reasoning Agent for Chest X-ray - ICML 2025
Go to https://mimic.physionet.org/ for access. Once you have the authority for the dataset, download the dataset at the data folder under the same directory as the this repository.
This repository presents a novel multi-agent conversation framework designed to enhance the capabilities of Large Language Models (LLMs) in diagnosing complex diseases. Our approach, structured under the Autogen framework, allows for in-depth conversations among LLMs, paving the way for more…
Anglin Liu 1, , Rundong Xue 2, , Xu R. Cao 3,† , Yifan Shen 3 , Yi Lu 1 , Xiang Li 3 , Qianqian Chen 4 , Jintai Chen 1,5,†
A Self-Evolving Multi-Agent Framework for Medical Multi-Disciplinary Team (MDT) Consultations.
A medical machine learning benchmark platform for evaluating automated machine learning agents on realistic healthcare tasks.
09/13/2024 🍓 We release new results and support for o1! - 08/17/2024 🎆 Major updates 🎇 - 🏥 A new suite of cases (AgentClinic-MIMIC-IV), based on real clinical cases from MIMIC-IV (requires approval from https://physionet.org/content/mimiciv/2.2/)! - More AgentClinic-MedQA cases 107 →…
DeepTumorVQA benchmark for VLMs and Agents (10k testing samples)
Observational studies can yield clinically actionable evidence at scale, but executing them on real-world databases is open-ended and requires coherent decisions across cohort construction, analysis, and reporting. Prior evaluations of LLM agents emphasize isolated steps or single answers,…
This benchmark system simulates an interactive conversation between a patient and an expert. The system evaluates how well participants' expert modules can handle realistic patient queries by either asking relevant questions or making final decisions based on the conversation history.
Patho-AgenticRAG : Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning
HealthFlow is a strategically self-evolving multi-agent framework for automating electronic health record (EHR) analysis. It turns a clinical analysis request into a governed workflow that plans, executes, evaluates, repairs, and writes back reusable experience for later tasks.
🎉 Our paper has been accepted to the NeurIPS 2025 Datasets & Benchmarks Track! 🎉
⭐ If you find this work helpful, please consider giving us a star! For questions and discussions, feel free to open an issue — we're happy to help.
A medagent for CDR selection and execution
C linical H ealthcare I n-Situ Environment Benchmark for long-horizon, policy-rich healthcare workflow agents
The code to visualize colon-bench (MICCAI 2026) data and run evaluations of MLLMs on the benchmark.
Healthcare Agent Orchestrator is a multi-agent accelerator that coordinates modular specialized agents across diverse data types and tools like M365 and Teams to assist multi-disciplinary healthcare workflows—such as cancer care.