Healthcare & Medical Index
Assesses model performance across the healthcare and medical domain. Capabilities evaluated include domain-specific knowledge (medicine, public health, biomedical sciences), clinical diagnosis and assessment, reasoning over long patient records and claims files, patient documentation, medication management, and more.
代表的なワークフローを見るThe Artificial Analysis Healthcare & Medical Index combines performance across benchmarks chosen for clinical and healthcare-support work, spanning medical knowledge, clinical reasoning, long-context reasoning over patient records, agentic workflows, and non-hallucination. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across healthcare tasks. 基礎となるすべてのベンチマークは、Artificial Analysisが独立して実施しています。評価方法の詳細は、知能ベンチマークの方法論をご覧ください。
| 能力 | 重み | 評価 |
|---|---|---|
| Medical & Health Knowledge | 30% | AA-Omniscience Health Accuracy |
| Agentic Knowledge Work | 25% | GDPval-AA v2、AA-Briefcase |
| Long-Context Reasoning | 15% | MLCR-AA |
| Non-Hallucination | 10% | AA-Omniscience Health Non-Hallucination |
| Reasoning | 10% | HLE |
| Agentic Tool Use | 10% | AutomationBench-AA Support & Operations |
スコア
Artificial Analysis Healthcare & Medical Index
Artificial Analysis Healthcare & Medical Index:能力の内訳
能力の内訳
Artificial Analysis Healthcare & Medical Index:Medical & Health Knowledge
代表的なワークフロー
Healthcare & Medical Indexで特に重視される能力を測る、実際の業務を想定したワークフローです。
例:Reassess a returning patient with worsening symptoms against the original EHR workup to build a differential from the new labs and imaging and surface alternative diagnoses the findings point to.
例:A surgical team that encounters unexpected anatomy mid-laparoscopic-procedure. Retrieve comparable case reports and imaging precedents and quickly output findings relevant to their immediate decision.
例:Turn a clinician's dictated notes from a follow-up visit into a structured SOAP note, pulling the patient's active problems and relevant history from the existing chart, placing each finding in the right section, and flagging the gaps the next provider would need filled.
例:Calculate a child's per-dose amount from their measurements and the prescriber's notes against the available suspension concentration, convert it to the millilitres to measure at each dose, and produce caregiver instructions that keep the total within the safe daily range.
例:Evaluate whether a dermatology team should adopt a newer procedure backed by emerging but limited long-term evidence to summarise the published trials and safety data, compare outcomes against the current standard of care, and outline the open questions the team still needs to resolve.
例:Work a several-hundred-page medical record assembled from multiple providers to reconstruct the treatment timeline, identify which encounters relate to the injury in question, and answer reviewer questions with citations to the underlying documents.
例:Turn a patient's after-visit summary into plain-language, step-by-step home-care instructions in their preferred language, anticipate the questions they are most likely to ask, and confirm the follow-up appointment and how to reach the clinic with concerns.
コスト
Artificial Analysis Healthcare & Medical Index:タスク当たりのコスト
Artificial Analysis Healthcare & Medical Index vs. タスク当たりのコスト
速度
Artificial Analysis Healthcare & Medical Index:タスク当たりの時間
出力トークン
Artificial Analysis Healthcare & Medical Index:タスク当たりの出力トークン
リリース日
Artificial Analysis Healthcare & Medical Index vs. リリース日
よくある質問
Artificial AnalysisのHealthcare & Medical Indexによると、現在医療・ヘルスケア業務で最も高い性能を示すAIモデルはClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (58)、Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (54)、Claude Opus 5 (Adaptive Reasoning, Max Effort) (53)です。ランキングは新しいモデルのリリースに合わせて更新されます。
はい。Artificial AnalysisのHealthcare & Medical Indexは、医療・ヘルスケア業務におけるAIモデルの性能を測る独立したベンチマークです。臨床知識、エージェント型ナレッジワーク、長い患者記録の推論、ハルシネーション抑制、臨床推論、エージェント型ツール使用を評価します。
Healthcare & Medical Indexは、医療・ヘルスケア分野におけるモデルの性能を評価するArtificial Analysisの複合ベンチマークです。医学、公衆衛生、生物医学などの専門知識に加え、臨床診断と評価、長い患者記録や保険請求資料の推論、患者記録作成、薬剤管理などを評価します。
Healthcare & Medical Indexは、各能力のサブスコアを加重平均して算出します。サブスコアとその重みは、Medical & Health Knowledge (30%)、Agentic Knowledge Work (25%)、Long-Context Reasoning (15%)、Non-Hallucination (10%)、Reasoning (10%)、Agentic Tool Use (10%)です。
Healthcare & Medical IndexにはAA-Omniscience Health Accuracy、GDPval-AA v2、AA-Briefcase、MLCR-AA、AA-Omniscience Health Non-Hallucination、HLE、AutomationBench-AA Support & Operationsが含まれます。
結果が公開されているモデルの中では、現在Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)がHealthcare & Medical Indexで最高の58を記録しています。 モデルを見る
Healthcare & Medical Indexのスコアが高いほど、指数を構成するベンチマーク全体での性能が優れていることを示します。特定の用途では、複合スコアよりも個別のベンチマーク結果のほうが参考になる場合があります。