医療・ヘルスケア指数
医療・ヘルスケア分野におけるモデルの性能を評価します。医学、公衆衛生、生物医学に関する専門知識、臨床診断・評価、長い患者記録や請求資料に対する推論、患者文書、服薬管理などを対象とします。
代表的なワークフローを見るThe Artificial Analysis Healthcare & Medical Index combines performance across benchmarks chosen for clinical and healthcare-support work, spanning medical knowledge, clinical reasoning, long-context reasoning over patient records, agentic workflows, and non-hallucination. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across healthcare tasks. 基礎となるすべてのベンチマークは、Artificial Analysisが独立して実施しています。評価方法の詳細は、知能ベンチマークの方法論をご覧ください。
| 能力 | 重み | 評価 |
|---|---|---|
| 医療・健康知識 | 30% | AA-Omniscience 健康の正確性 |
| エージェント型知識作業 | 25% | GDPval-AA v2.1、AA-Briefcase v1.1 |
| 長文脈推論 | 15% | MLCR-AA |
| 非ハルシネーション | 10% | AA-Omniscience 健康の非ハルシネーション |
| 推論 | 10% | HLE |
| エージェント型ツール利用 | 10% | AutomationBench-AA サポート、オペレーション |
スコア
Artificial Analysis 医療・ヘルスケア指数
Artificial Analysis 医療・ヘルスケア指数:能力の内訳
能力の内訳
Artificial Analysis 医療・ヘルスケア指数:医療・健康知識
代表的なワークフロー
医療・ヘルスケア指数で特に重視される能力を測る、実際の業務を想定したワークフローです。
例:Reassess a returning patient with worsening symptoms against the original EHR workup to build a differential from the new labs and imaging and surface alternative diagnoses the findings point to.
例:A surgical team that encounters unexpected anatomy mid-laparoscopic-procedure. Retrieve comparable case reports and imaging precedents and quickly output findings relevant to their immediate decision.
例:Turn a clinician's dictated notes from a follow-up visit into a structured SOAP note, pulling the patient's active problems and relevant history from the existing chart, placing each finding in the right section, and flagging the gaps the next provider would need filled.
例:Calculate a child's per-dose amount from their measurements and the prescriber's notes against the available suspension concentration, convert it to the millilitres to measure at each dose, and produce caregiver instructions that keep the total within the safe daily range.
例:Evaluate whether a dermatology team should adopt a newer procedure backed by emerging but limited long-term evidence to summarise the published trials and safety data, compare outcomes against the current standard of care, and outline the open questions the team still needs to resolve.
例:Work a several-hundred-page medical record assembled from multiple providers to reconstruct the treatment timeline, identify which encounters relate to the injury in question, and answer reviewer questions with citations to the underlying documents.
例:Turn a patient's after-visit summary into plain-language, step-by-step home-care instructions in their preferred language, anticipate the questions they are most likely to ask, and confirm the follow-up appointment and how to reach the clinic with concerns.
コスト
Artificial Analysis 医療・ヘルスケア指数:タスク当たりのコスト
Artificial Analysis 医療・ヘルスケア指数 vs. タスク当たりのコスト
速度
Artificial Analysis 医療・ヘルスケア指数:タスク当たりの時間
出力トークン
Artificial Analysis 医療・ヘルスケア指数:タスク当たりの出力トークン
リリース日
Artificial Analysis 医療・ヘルスケア指数 vs. リリース日
よくある質問
Artificial Analysisの医療・ヘルスケア指数によると、現在医療・ヘルスケア業務で最も高い性能を示すAIモデルはClaude Opus 5.5 (Max, Default Fallback) (61)、Claude Sonnet 5.5 (Max, Default Fallback) (58)、Claude Fable 5.1 (Max, Default Fallback) (58)です。ランキングは新しいモデルのリリースに合わせて更新されます。
はい。Artificial Analysisの医療・ヘルスケア指数は、医療・ヘルスケア業務におけるAIモデルの性能を測る独立したベンチマークです。臨床知識、エージェント型ナレッジワーク、長い患者記録の推論、ハルシネーション抑制、臨床推論、エージェント型ツール使用を評価します。
医療・ヘルスケア指数は、医療・ヘルスケア分野におけるモデルの性能を評価するArtificial Analysisの複合ベンチマークです。医学、公衆衛生、生物医学などの専門知識に加え、臨床診断と評価、長い患者記録や保険請求資料の推論、患者記録作成、薬剤管理などを評価します。
医療・ヘルスケア指数は、各能力のサブスコアを加重平均して算出します。サブスコアとその重みは、医療・健康知識 (30%)、エージェント型知識作業 (25%)、長文脈推論 (15%)、非ハルシネーション (10%)、推論 (10%)、エージェント型ツール利用 (10%)です。
医療・ヘルスケア指数にはAA-Omniscience 健康の正確性、GDPval-AA v2.1、AA-Briefcase v1.1、MLCR-AA、AA-Omniscience 健康の非ハルシネーション、HLE、AutomationBench-AA サポート、オペレーションが含まれます。
結果が公開されているモデルの中では、現在Claude Opus 5.5 (Max, Default Fallback)が医療・ヘルスケア指数で最高の61を記録しています。 モデルを見る
医療・ヘルスケア指数のスコアが高いほど、指数を構成するベンチマーク全体での性能が優れていることを示します。特定の用途では、複合スコアよりも個別のベンチマーク結果のほうが参考になる場合があります。