医疗与健康指数
评估模型在医疗与健康领域的表现。评估的能力包括医学、公共卫生、生物医学等领域知识,以及临床诊断与评估、对长篇患者记录和理赔文件的推理、患者文档、药物管理等。
查看代表性工作流The Artificial Analysis Healthcare & Medical Index combines performance across benchmarks chosen for clinical and healthcare-support work, spanning medical knowledge, clinical reasoning, long-context reasoning over patient records, agentic workflows, and non-hallucination. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across healthcare tasks. 所有底层基准测试均由 Artificial Analysis 独立运行。有关评测的实施方式,请参阅我们的智能基准测试方法论。
| 能力 | 权重 | 评测 |
|---|---|---|
| 医疗与健康知识 | 30% | AA-Omniscience 健康准确率 |
| 智能体知识工作 | 25% | GDPval-AA v2.1和AA-Briefcase v1.1 |
| 长上下文推理 | 15% | MLCR-AA |
| 避免幻觉 | 10% | AA-Omniscience 健康避免幻觉 |
| 推理 | 10% | HLE |
| 智能体工具使用 | 10% | AutomationBench-AA 支持和运营 |
得分
Artificial Analysis 医疗与健康指数
Artificial Analysis 医疗与健康指数:能力明细
能力明细
Artificial Analysis 医疗与健康指数:医疗与健康知识
代表性工作流
这些真实工作流重点检验 医疗与健康指数 中权重最高的能力。
示例:Reassess a returning patient with worsening symptoms against the original EHR workup to build a differential from the new labs and imaging and surface alternative diagnoses the findings point to.
示例:A surgical team that encounters unexpected anatomy mid-laparoscopic-procedure. Retrieve comparable case reports and imaging precedents and quickly output findings relevant to their immediate decision.
示例:Turn a clinician's dictated notes from a follow-up visit into a structured SOAP note, pulling the patient's active problems and relevant history from the existing chart, placing each finding in the right section, and flagging the gaps the next provider would need filled.
示例:Calculate a child's per-dose amount from their measurements and the prescriber's notes against the available suspension concentration, convert it to the millilitres to measure at each dose, and produce caregiver instructions that keep the total within the safe daily range.
示例:Evaluate whether a dermatology team should adopt a newer procedure backed by emerging but limited long-term evidence to summarise the published trials and safety data, compare outcomes against the current standard of care, and outline the open questions the team still needs to resolve.
示例:Work a several-hundred-page medical record assembled from multiple providers to reconstruct the treatment timeline, identify which encounters relate to the injury in question, and answer reviewer questions with citations to the underlying documents.
示例:Turn a patient's after-visit summary into plain-language, step-by-step home-care instructions in their preferred language, anticipate the questions they are most likely to ask, and confirm the follow-up appointment and how to reach the clinic with concerns.
成本
Artificial Analysis 医疗与健康指数:单任务成本
Artificial Analysis 医疗与健康指数 与单任务成本
速度
Artificial Analysis 医疗与健康指数:单任务耗时
输出 token
Artificial Analysis 医疗与健康指数:单任务输出 token
发布日期
Artificial Analysis 医疗与健康指数 与发布日期
常见问题
根据 Artificial Analysis 医疗与健康指数,目前在医疗健康工作上表现最佳的 AI 模型是 Claude Opus 5.5 (Max, Default Fallback) (61)、Claude Sonnet 5.5 (Max, Default Fallback) (58)和Claude Fable 5.1 (Max, Default Fallback) (58)。新模型发布后,排行榜会随之更新。
有。Artificial Analysis 医疗与健康指数 是一项独立基准测试,用于衡量 AI 模型在医疗健康工作上的表现。它评估临床知识、智能体知识工作、长篇患者记录推理、避免幻觉、临床推理和智能体工具使用等能力。
医疗与健康指数 是 Artificial Analysis 推出的综合基准测试,用于评估模型在医疗健康领域的表现。评估能力包括医学、公共卫生、生物医学科学等专业知识,以及临床诊断与评估、长篇患者记录和理赔文件推理、患者文档记录、用药管理等。
医疗与健康指数 按各项能力子分数的加权平均值计算。各项子分数及其权重为:医疗与健康知识 (30%)、智能体知识工作 (25%)、长上下文推理 (15%)、避免幻觉 (10%)、推理 (10%)和智能体工具使用 (10%)。
医疗与健康指数 包含 AA-Omniscience 健康准确率、GDPval-AA v2.1、AA-Briefcase v1.1、MLCR-AA、AA-Omniscience 健康避免幻觉、HLE和AutomationBench-AA 支持和运营。
在已公布结果的模型中,Claude Opus 5.5 (Max, Default Fallback) 目前以 61 分位居 医疗与健康指数 榜首。 查看模型
医疗与健康指数 得分越高,表示模型在构成该指数的各项基准测试中整体表现越强。对于特定用例,单项基准测试结果可能比综合得分更具参考价值。