すべての能力指数

Healthcare & Medical Index

Assesses model performance across the healthcare and medical domain. Capabilities evaluated include domain-specific knowledge (medicine, public health, biomedical sciences), clinical diagnosis and assessment, reasoning over long patient records and claims files, patient documentation, medication management, and more.

代表的なワークフローを見る

スコア

Artificial Analysis Healthcare & Medical Index

Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2, AA-Briefcase, MLCR-AA, Humanity's Last Exam, AutomationBench-AA · Higher is better

Artificial Analysis Healthcare & Medical Index:能力の内訳

Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2, AA-Briefcase, MLCR-AA, Humanity's Last Exam, AutomationBench-AA · 貢献度別に分割

能力の内訳

Artificial Analysis Healthcare & Medical Index:Medical & Health Knowledge

Incorporates 1 evaluation: AA-Omniscience · Higher is better

代表的なワークフロー

Healthcare & Medical Indexで特に重視される能力を測る、実際の業務を想定したワークフローです。

Patient diagnosis & clinical assessmentMedical & Health KnowledgeLong-Context ReasoningNon-HallucinationReasoning

例:Reassess a returning patient with worsening symptoms against the original EHR workup to build a differential from the new labs and imaging and surface alternative diagnoses the findings point to.

Treatment administration & proceduresMedical & Health KnowledgeNon-Hallucination

例:A surgical team that encounters unexpected anatomy mid-laparoscopic-procedure. Retrieve comparable case reports and imaging precedents and quickly output findings relevant to their immediate decision.

Patient documentation & chartingMedical & Health KnowledgeAgentic Knowledge Work

例:Turn a clinician's dictated notes from a follow-up visit into a structured SOAP note, pulling the patient's active problems and relevant history from the existing chart, placing each finding in the right section, and flagging the gaps the next provider would need filled.

Medication management & pharmacy coordinationMedical & Health KnowledgeNon-HallucinationReasoning

例:Calculate a child's per-dose amount from their measurements and the prescriber's notes against the available suspension concentration, convert it to the millilitres to measure at each dose, and produce caregiver instructions that keep the total within the safe daily range.

Continuing education & researchMedical & Health KnowledgeAgentic Knowledge WorkLong-Context ReasoningReasoning

例:Evaluate whether a dermatology team should adopt a newer procedure backed by emerging but limited long-term evidence to summarise the published trials and safety data, compare outcomes against the current standard of care, and outline the open questions the team still needs to resolve.

Medical records review & claimsLong-Context ReasoningNon-HallucinationReasoning

例:Work a several-hundred-page medical record assembled from multiple providers to reconstruct the treatment timeline, identify which encounters relate to the injury in question, and answer reviewer questions with citations to the underlying documents.

Patient education & care coordinationMedical & Health KnowledgeAgentic Knowledge WorkAgentic Tool Use

例:Turn a patient's after-visit summary into plain-language, step-by-step home-care instructions in their preferred language, anticipate the questions they are most likely to ask, and confirm the follow-up appointment and how to reach the clinic with concerns.

コスト

Artificial Analysis Healthcare & Medical Index:タスク当たりのコスト

タスク当たりの平均コスト(USD)を、入力、キャッシュヒット、キャッシュ書き込み、推論、回答の各トークンに分けて表示

Artificial Analysis Healthcare & Medical Index vs. タスク当たりのコスト

Artificial Analysis Healthcare & Medical Index · タスク当たりの加重平均コスト(USD)
Most attractive quadrant
Pareto line

速度

Artificial Analysis Healthcare & Medical Index:タスク当たりの時間

Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is better

出力トークン

Artificial Analysis Healthcare & Medical Index:タスク当たりの出力トークン

1件のタスクの実行に使用した出力トークンを、推論トークンと回答トークンに分けて表示

リリース日

Artificial Analysis Healthcare & Medical Index vs. リリース日

Most attractive region

よくある質問