Healthcare & Medical Index
Assesses model performance across the healthcare and medical domain. Capabilities evaluated include domain-specific knowledge (medicine, public health, biomedical sciences), clinical diagnosis and assessment, reasoning over long patient records and claims files, patient documentation, medication management, and more.
대표 워크플로 보기The Artificial Analysis Healthcare & Medical Index combines performance across benchmarks chosen for clinical and healthcare-support work, spanning medical knowledge, clinical reasoning, long-context reasoning over patient records, agentic workflows, and non-hallucination. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across healthcare tasks. 모든 기반 벤치마크는 Artificial Analysis가 독립적으로 실행합니다. 평가 수행 방식은 지능 벤치마킹 방법론을 참조하세요.
| 역량 | 가중치 | 평가 |
|---|---|---|
| Medical & Health Knowledge | 30% | AA-Omniscience Health Accuracy |
| Agentic Knowledge Work | 25% | GDPval-AA v2 및 AA-Briefcase |
| Long-Context Reasoning | 15% | MLCR-AA |
| Non-Hallucination | 10% | AA-Omniscience Health Non-Hallucination |
| Reasoning | 10% | HLE |
| Agentic Tool Use | 10% | AutomationBench-AA Support & Operations |
점수
Artificial Analysis Healthcare & Medical Index
Artificial Analysis Healthcare & Medical Index: 역량 세부 분석
역량 세부 분석
Artificial Analysis Healthcare & Medical Index: Medical & Health Knowledge
대표 워크플로
Healthcare & Medical Index에서 가장 높은 가중치를 둔 역량을 평가하는 실제 워크플로입니다.
예시: Reassess a returning patient with worsening symptoms against the original EHR workup to build a differential from the new labs and imaging and surface alternative diagnoses the findings point to.
예시: A surgical team that encounters unexpected anatomy mid-laparoscopic-procedure. Retrieve comparable case reports and imaging precedents and quickly output findings relevant to their immediate decision.
예시: Turn a clinician's dictated notes from a follow-up visit into a structured SOAP note, pulling the patient's active problems and relevant history from the existing chart, placing each finding in the right section, and flagging the gaps the next provider would need filled.
예시: Calculate a child's per-dose amount from their measurements and the prescriber's notes against the available suspension concentration, convert it to the millilitres to measure at each dose, and produce caregiver instructions that keep the total within the safe daily range.
예시: Evaluate whether a dermatology team should adopt a newer procedure backed by emerging but limited long-term evidence to summarise the published trials and safety data, compare outcomes against the current standard of care, and outline the open questions the team still needs to resolve.
예시: Work a several-hundred-page medical record assembled from multiple providers to reconstruct the treatment timeline, identify which encounters relate to the injury in question, and answer reviewer questions with citations to the underlying documents.
예시: Turn a patient's after-visit summary into plain-language, step-by-step home-care instructions in their preferred language, anticipate the questions they are most likely to ask, and confirm the follow-up appointment and how to reach the clinic with concerns.
비용
Artificial Analysis Healthcare & Medical Index: 작업당 비용
Artificial Analysis Healthcare & Medical Index vs. 작업당 비용
속도
Artificial Analysis Healthcare & Medical Index: 작업당 시간
출력 토큰
Artificial Analysis Healthcare & Medical Index: 작업당 출력 토큰
출시일
Artificial Analysis Healthcare & Medical Index vs. 출시일
자주 묻는 질문
Artificial Analysis Healthcare & Medical Index에 따르면 현재 의료 및 헬스케어 업무에서 가장 뛰어난 AI 모델은 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (58), Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (54) 및 Claude Opus 5 (Adaptive Reasoning, Max Effort) (53)입니다. 순위는 새 모델 출시에 맞춰 업데이트됩니다.
네. Artificial Analysis의 Healthcare & Medical Index는 의료 및 헬스케어 업무에서 AI 모델의 성능을 측정하는 독립적인 벤치마크입니다. 임상 지식, 에이전트 지식 업무, 긴 환자 기록 추론, 환각 방지, 임상 추론, 에이전트 도구 사용을 평가합니다.
Healthcare & Medical Index는 의료 및 헬스케어 분야에서 모델의 성능을 평가하는 Artificial Analysis의 종합 벤치마크입니다. 의학, 공중 보건, 생의학 관련 전문 지식과 임상 진단 및 평가, 긴 환자 기록과 보험 청구 파일에 대한 추론, 환자 문서화, 약물 관리 등을 평가합니다.
Healthcare & Medical Index는 각 역량 하위 점수의 가중 평균으로 계산합니다. 하위 점수와 가중치는 Medical & Health Knowledge (30%), Agentic Knowledge Work (25%), Long-Context Reasoning (15%), Non-Hallucination (10%), Reasoning (10%) 및 Agentic Tool Use (10%)입니다.
Healthcare & Medical Index에는 AA-Omniscience Health Accuracy, GDPval-AA v2, AA-Briefcase, MLCR-AA, AA-Omniscience Health Non-Hallucination, HLE 및 AutomationBench-AA Support & Operations이 포함됩니다.
결과가 공개된 모델 중 현재 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)의 Healthcare & Medical Index 점수가 58로 가장 높습니다. 모델 보기
Healthcare & Medical Index 점수가 높을수록 지수를 구성하는 벤치마크 전반에서 더 뛰어난 성능을 보인다는 뜻입니다. 특정 사용 사례에서는 종합 점수보다 개별 벤치마크 결과가 더 유용할 수 있습니다.