Healthcare & Medical Index
Assesses model performance across the healthcare and medical domain. Capabilities evaluated include domain-specific knowledge (medicine, public health, biomedical sciences), clinical diagnosis and assessment, reasoning over long patient records and claims files, patient documentation, medication management, and more.
Repräsentative Workflows ansehenThe Artificial Analysis Healthcare & Medical Index combines performance across benchmarks chosen for clinical and healthcare-support work, spanning medical knowledge, clinical reasoning, long-context reasoning over patient records, agentic workflows, and non-hallucination. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across healthcare tasks. Alle zugrunde liegenden Benchmarks werden unabhängig von Artificial Analysis durchgeführt. Wie die Evaluierungen ablaufen, erläutert unsere Methodik für das Intelligenz-Benchmarking.
| Fähigkeit | Gewichtung | Evaluierungen |
|---|---|---|
| Medical & Health Knowledge | 30 % | AA-Omniscience Health Accuracy |
| Agentic Knowledge Work | 25 % | GDPval-AA v2 und AA-Briefcase |
| Long-Context Reasoning | 15 % | MLCR-AA |
| Non-Hallucination | 10 % | AA-Omniscience Health Non-Hallucination |
| Reasoning | 10 % | HLE |
| Agentic Tool Use | 10 % | AutomationBench-AA Support & Operations |
Punktzahl
Artificial Analysis Healthcare & Medical Index
Artificial Analysis Healthcare & Medical Index: Aufschlüsselung der Fähigkeiten
Aufschlüsselung der Fähigkeiten
Artificial Analysis Healthcare & Medical Index: Medical & Health Knowledge
Repräsentative Workflows
Praxisnahe Workflows, die besonders die von Healthcare & Medical Index am stärksten gewichteten Fähigkeiten prüfen.
Beispiel: Reassess a returning patient with worsening symptoms against the original EHR workup to build a differential from the new labs and imaging and surface alternative diagnoses the findings point to.
Beispiel: A surgical team that encounters unexpected anatomy mid-laparoscopic-procedure. Retrieve comparable case reports and imaging precedents and quickly output findings relevant to their immediate decision.
Beispiel: Turn a clinician's dictated notes from a follow-up visit into a structured SOAP note, pulling the patient's active problems and relevant history from the existing chart, placing each finding in the right section, and flagging the gaps the next provider would need filled.
Beispiel: Calculate a child's per-dose amount from their measurements and the prescriber's notes against the available suspension concentration, convert it to the millilitres to measure at each dose, and produce caregiver instructions that keep the total within the safe daily range.
Beispiel: Evaluate whether a dermatology team should adopt a newer procedure backed by emerging but limited long-term evidence to summarise the published trials and safety data, compare outcomes against the current standard of care, and outline the open questions the team still needs to resolve.
Beispiel: Work a several-hundred-page medical record assembled from multiple providers to reconstruct the treatment timeline, identify which encounters relate to the injury in question, and answer reviewer questions with citations to the underlying documents.
Beispiel: Turn a patient's after-visit summary into plain-language, step-by-step home-care instructions in their preferred language, anticipate the questions they are most likely to ask, and confirm the follow-up appointment and how to reach the clinic with concerns.
Kosten
Artificial Analysis Healthcare & Medical Index: Kosten pro Aufgabe
Artificial Analysis Healthcare & Medical Index vs. Kosten pro Aufgabe
Geschwindigkeit
Artificial Analysis Healthcare & Medical Index: Zeit pro Aufgabe
Ausgabe-Token
Artificial Analysis Healthcare & Medical Index: Ausgabe-Token pro Aufgabe
Veröffentlichungsdatum
Artificial Analysis Healthcare & Medical Index vs. Veröffentlichungsdatum
Häufig gestellte Fragen
Laut dem Artificial Analysis Healthcare & Medical Index sind Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (58), Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (54) und Claude Opus 5 (Adaptive Reasoning, Max Effort) (53) derzeit die leistungsstärksten KI-Modelle für Aufgaben in Gesundheit und Medizin. Die Rangliste wird bei der Veröffentlichung neuer Modelle aktualisiert.
Ja. Der Healthcare & Medical Index von Artificial Analysis ist ein unabhängiger Benchmark für die Leistung von KI-Modellen bei Aufgaben in Gesundheit und Medizin. Er misst klinisches Wissen, agentische Wissensarbeit, Schlussfolgern über lange Patientenakten, Halluzinationsfreiheit, klinisches Schlussfolgern und agentische Tool-Nutzung.
Der Healthcare & Medical Index ist ein zusammengesetzter Benchmark von Artificial Analysis, der die Modellleistung in Gesundheit und Medizin bewertet. Geprüft werden unter anderem Fachwissen zu Medizin, öffentlicher Gesundheit und Biomedizin, klinische Diagnose und Beurteilung, Schlussfolgern über lange Patienten- und Leistungsakten, Patientendokumentation und Medikamentenmanagement.
Der Healthcare & Medical Index wird als gewichteter Durchschnitt seiner Teilpunktzahlen berechnet. Die Teilpunktzahlen und ihre Gewichtungen sind: Medical & Health Knowledge (30 %), Agentic Knowledge Work (25 %), Long-Context Reasoning (15 %), Non-Hallucination (10 %), Reasoning (10 %) und Agentic Tool Use (10 %).
Der Healthcare & Medical Index enthält AA-Omniscience Health Accuracy, GDPval-AA v2, AA-Briefcase, MLCR-AA, AA-Omniscience Health Non-Hallucination, HLE und AutomationBench-AA Support & Operations.
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) erzielt derzeit mit 58 die höchste Punktzahl im Healthcare & Medical Index unter den Modellen mit veröffentlichten Ergebnissen. Modell ansehen
Eine höhere Punktzahl im Healthcare & Medical Index steht für eine insgesamt stärkere Leistung in den Benchmarks des Index. Für einen bestimmten Anwendungsfall können einzelne Benchmark-Ergebnisse aussagekräftiger sein als die zusammengesetzte Punktzahl.