Healthcare & Medical Index
Assesses model performance across the healthcare and medical domain. Capabilities evaluated include domain-specific knowledge (medicine, public health, biomedical sciences), clinical diagnosis and assessment, reasoning over long patient records and claims files, patient documentation, medication management, and more.
See representative workflowsThe Artificial Analysis Healthcare & Medical Index combines performance across benchmarks chosen for clinical and healthcare-support work, spanning medical knowledge, clinical reasoning, long-context reasoning over patient records, agentic workflows, and non-hallucination. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across healthcare tasks. All underlying benchmarks are run independently by Artificial Analysis. See our Intelligence Benchmarking Methodology for how evaluations are conducted.
| Capability | Weight | Evaluations |
|---|---|---|
| Medical & Health Knowledge | 30% | AA-Omniscience Health Accuracy |
| Agentic Knowledge Work | 25% | GDPval-AA v2 and AA-Briefcase |
| Long-Context Reasoning | 15% | MLCR-AA |
| Non-Hallucination | 10% | AA-Omniscience Health Non-Hallucination |
| Reasoning | 10% | HLE |
| Agentic Tool Use | 10% | AutomationBench-AA Support & Operations |
Score
Artificial Analysis Healthcare & Medical Index
Artificial Analysis Healthcare & Medical Index: Capability Breakdown
Capability Breakdown
Artificial Analysis Healthcare & Medical Index: Medical & Health Knowledge
Representative Workflows
Real-world workflows that exercise the capabilities the Healthcare & Medical Index weights most heavily.
Example: Reassess a returning patient with worsening symptoms against the original EHR workup to build a differential from the new labs and imaging and surface alternative diagnoses the findings point to.
Example: A surgical team that encounters unexpected anatomy mid-laparoscopic-procedure. Retrieve comparable case reports and imaging precedents and quickly output findings relevant to their immediate decision.
Example: Turn a clinician's dictated notes from a follow-up visit into a structured SOAP note, pulling the patient's active problems and relevant history from the existing chart, placing each finding in the right section, and flagging the gaps the next provider would need filled.
Example: Calculate a child's per-dose amount from their measurements and the prescriber's notes against the available suspension concentration, convert it to the millilitres to measure at each dose, and produce caregiver instructions that keep the total within the safe daily range.
Example: Evaluate whether a dermatology team should adopt a newer procedure backed by emerging but limited long-term evidence to summarise the published trials and safety data, compare outcomes against the current standard of care, and outline the open questions the team still needs to resolve.
Example: Work a several-hundred-page medical record assembled from multiple providers to reconstruct the treatment timeline, identify which encounters relate to the injury in question, and answer reviewer questions with citations to the underlying documents.
Example: Turn a patient's after-visit summary into plain-language, step-by-step home-care instructions in their preferred language, anticipate the questions they are most likely to ask, and confirm the follow-up appointment and how to reach the clinic with concerns.
Cost
Artificial Analysis Healthcare & Medical Index: Cost per Task
Artificial Analysis Healthcare & Medical Index vs. Cost per Task
Speed
Artificial Analysis Healthcare & Medical Index: Time per Task
Output Tokens
Artificial Analysis Healthcare & Medical Index: Output Tokens per Task
Release Date
Artificial Analysis Healthcare & Medical Index vs. Release Date
Frequently Asked Questions
Based on the Artificial Analysis Healthcare & Medical Index, the top-performing AI models for healthcare and medical work are currently Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (58), Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (54), and Claude Opus 5 (Adaptive Reasoning, Max Effort) (53). Rankings are updated as new models are released.
Yes. The Healthcare & Medical Index from Artificial Analysis is an independent benchmark of how AI models perform on healthcare and medical work. It measures performance on the medical and healthcare domain, including clinical knowledge, agentic knowledge work, long-context reasoning over patient records, non-hallucination, clinical reasoning, and agentic tool use.
The Healthcare & Medical Index is a composite benchmark from Artificial Analysis that assesses model performance across the healthcare and medical domain. Capabilities evaluated include domain-specific knowledge (medicine, public health, biomedical sciences), clinical diagnosis and assessment, reasoning over long patient records and claims files, patient documentation, medication management, and more.
The Healthcare & Medical Index is calculated as a weighted average of its capability sub-scores. The sub-scores and their weights are: Medical & Health Knowledge (30%), Agentic Knowledge Work (25%), Long-Context Reasoning (15%), Non-Hallucination (10%), Reasoning (10%), and Agentic Tool Use (10%).
The Healthcare & Medical Index includes AA-Omniscience Health Accuracy, GDPval-AA v2, AA-Briefcase, MLCR-AA, AA-Omniscience Health Non-Hallucination, HLE, and AutomationBench-AA Support & Operations.
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) currently has the highest Healthcare & Medical Index score, with a score of 58 among models with published results. View model
A higher Healthcare & Medical Index score indicates stronger overall performance across the benchmarks that make up the index. For a specific use case, individual benchmark results may be more informative than the composite score.