Healthcare & Medical Index
Assesses model performance across the healthcare and medical domain. Capabilities evaluated include domain-specific knowledge (medicine, public health, biomedical sciences), clinical diagnosis and assessment, reasoning over long patient records and claims files, patient documentation, medication management, and more.
查看代表性工作流The Artificial Analysis Healthcare & Medical Index combines performance across benchmarks chosen for clinical and healthcare-support work, spanning medical knowledge, clinical reasoning, long-context reasoning over patient records, agentic workflows, and non-hallucination. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across healthcare tasks. 所有底层基准测试均由 Artificial Analysis 独立运行。有关评测的实施方式,请参阅我们的智能基准测试方法论。
| 能力 | 权重 | 评测 |
|---|---|---|
| Medical & Health Knowledge | 30% | AA-Omniscience Health Accuracy |
| Agentic Knowledge Work | 25% | GDPval-AA v2和AA-Briefcase |
| Long-Context Reasoning | 15% | MLCR-AA |
| Non-Hallucination | 10% | AA-Omniscience Health Non-Hallucination |
| Reasoning | 10% | HLE |
| Agentic Tool Use | 10% | AutomationBench-AA Support & Operations |
得分
Artificial Analysis Healthcare & Medical Index
Artificial Analysis Healthcare & Medical Index:能力明细
能力明细
Artificial Analysis Healthcare & Medical Index:Medical & Health Knowledge
代表性工作流
这些真实工作流重点检验 Healthcare & Medical Index 中权重最高的能力。
示例:Reassess a returning patient with worsening symptoms against the original EHR workup to build a differential from the new labs and imaging and surface alternative diagnoses the findings point to.
示例:A surgical team that encounters unexpected anatomy mid-laparoscopic-procedure. Retrieve comparable case reports and imaging precedents and quickly output findings relevant to their immediate decision.
示例:Turn a clinician's dictated notes from a follow-up visit into a structured SOAP note, pulling the patient's active problems and relevant history from the existing chart, placing each finding in the right section, and flagging the gaps the next provider would need filled.
示例:Calculate a child's per-dose amount from their measurements and the prescriber's notes against the available suspension concentration, convert it to the millilitres to measure at each dose, and produce caregiver instructions that keep the total within the safe daily range.
示例:Evaluate whether a dermatology team should adopt a newer procedure backed by emerging but limited long-term evidence to summarise the published trials and safety data, compare outcomes against the current standard of care, and outline the open questions the team still needs to resolve.
示例:Work a several-hundred-page medical record assembled from multiple providers to reconstruct the treatment timeline, identify which encounters relate to the injury in question, and answer reviewer questions with citations to the underlying documents.
示例:Turn a patient's after-visit summary into plain-language, step-by-step home-care instructions in their preferred language, anticipate the questions they are most likely to ask, and confirm the follow-up appointment and how to reach the clinic with concerns.
成本
Artificial Analysis Healthcare & Medical Index:单任务成本
Artificial Analysis Healthcare & Medical Index 与单任务成本
速度
Artificial Analysis Healthcare & Medical Index:单任务耗时
输出 token
Artificial Analysis Healthcare & Medical Index:单任务输出 token
发布日期
Artificial Analysis Healthcare & Medical Index 与发布日期
常见问题
根据 Artificial Analysis Healthcare & Medical Index,目前在医疗健康工作上表现最佳的 AI 模型是 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (58)、Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (54)和Claude Opus 5 (Adaptive Reasoning, Max Effort) (53)。新模型发布后,排行榜会随之更新。
有。Artificial Analysis Healthcare & Medical Index 是一项独立基准测试,用于衡量 AI 模型在医疗健康工作上的表现。它评估临床知识、智能体知识工作、长篇患者记录推理、避免幻觉、临床推理和智能体工具使用等能力。
Healthcare & Medical Index 是 Artificial Analysis 推出的综合基准测试,用于评估模型在医疗健康领域的表现。评估能力包括医学、公共卫生、生物医学科学等专业知识,以及临床诊断与评估、长篇患者记录和理赔文件推理、患者文档记录、用药管理等。
Healthcare & Medical Index 按各项能力子分数的加权平均值计算。各项子分数及其权重为:Medical & Health Knowledge (30%)、Agentic Knowledge Work (25%)、Long-Context Reasoning (15%)、Non-Hallucination (10%)、Reasoning (10%)和Agentic Tool Use (10%)。
Healthcare & Medical Index 包含 AA-Omniscience Health Accuracy、GDPval-AA v2、AA-Briefcase、MLCR-AA、AA-Omniscience Health Non-Hallucination、HLE和AutomationBench-AA Support & Operations。
在已公布结果的模型中,Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) 目前以 58 分位居 Healthcare & Medical Index 榜首。 查看模型
Healthcare & Medical Index 得分越高,表示模型在构成该指数的各项基准测试中整体表现越强。对于特定用例,单项基准测试结果可能比综合得分更具参考价值。