의료 및 헬스케어 지수
의료 및 헬스케어 전반에서 모델 성능을 평가합니다. 의학, 공중보건, 생의학에 관한 전문 지식, 임상 진단 및 평가, 긴 환자 기록과 청구 파일에 대한 추론, 환자 문서화, 의약품 관리 등을 평가합니다.
대표 워크플로 보기The Artificial Analysis Healthcare & Medical Index combines performance across benchmarks chosen for clinical and healthcare-support work, spanning medical knowledge, clinical reasoning, long-context reasoning over patient records, agentic workflows, and non-hallucination. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across healthcare tasks. 모든 기반 벤치마크는 Artificial Analysis가 독립적으로 실행합니다. 평가 수행 방식은 지능 벤치마킹 방법론을 참조하세요.
| 역량 | 가중치 | 평가 |
|---|---|---|
| 의료 및 건강 지식 | 30% | AA-Omniscience 건강 정확도 |
| 에이전트형 지식 작업 | 25% | GDPval-AA v2.1 및 AA-Briefcase v1.1 |
| 장문 맥락 추론 | 15% | MLCR-AA |
| 환각 방지 | 10% | AA-Omniscience 건강 환각 방지 |
| 추론 | 10% | HLE |
| 에이전트형 도구 사용 | 10% | AutomationBench-AA 지원 및 운영 |
점수
Artificial Analysis 의료 및 헬스케어 지수
Artificial Analysis 의료 및 헬스케어 지수: 역량 세부 분석
역량 세부 분석
Artificial Analysis 의료 및 헬스케어 지수: 의료 및 건강 지식
대표 워크플로
의료 및 헬스케어 지수에서 가장 높은 가중치를 둔 역량을 평가하는 실제 워크플로입니다.
예시: Reassess a returning patient with worsening symptoms against the original EHR workup to build a differential from the new labs and imaging and surface alternative diagnoses the findings point to.
예시: A surgical team that encounters unexpected anatomy mid-laparoscopic-procedure. Retrieve comparable case reports and imaging precedents and quickly output findings relevant to their immediate decision.
예시: Turn a clinician's dictated notes from a follow-up visit into a structured SOAP note, pulling the patient's active problems and relevant history from the existing chart, placing each finding in the right section, and flagging the gaps the next provider would need filled.
예시: Calculate a child's per-dose amount from their measurements and the prescriber's notes against the available suspension concentration, convert it to the millilitres to measure at each dose, and produce caregiver instructions that keep the total within the safe daily range.
예시: Evaluate whether a dermatology team should adopt a newer procedure backed by emerging but limited long-term evidence to summarise the published trials and safety data, compare outcomes against the current standard of care, and outline the open questions the team still needs to resolve.
예시: Work a several-hundred-page medical record assembled from multiple providers to reconstruct the treatment timeline, identify which encounters relate to the injury in question, and answer reviewer questions with citations to the underlying documents.
예시: Turn a patient's after-visit summary into plain-language, step-by-step home-care instructions in their preferred language, anticipate the questions they are most likely to ask, and confirm the follow-up appointment and how to reach the clinic with concerns.
비용
Artificial Analysis 의료 및 헬스케어 지수: 작업당 비용
Artificial Analysis 의료 및 헬스케어 지수 vs. 작업당 비용
속도
Artificial Analysis 의료 및 헬스케어 지수: 작업당 시간
출력 토큰
Artificial Analysis 의료 및 헬스케어 지수: 작업당 출력 토큰
출시일
Artificial Analysis 의료 및 헬스케어 지수 vs. 출시일
자주 묻는 질문
Artificial Analysis 의료 및 헬스케어 지수에 따르면 현재 의료 및 헬스케어 업무에서 가장 뛰어난 AI 모델은 Claude Opus 5.5 (Max, Default Fallback) (61), Claude Sonnet 5.5 (Max, Default Fallback) (58) 및 Claude Fable 5.1 (Max, Default Fallback) (58)입니다. 순위는 새 모델 출시에 맞춰 업데이트됩니다.
네. Artificial Analysis의 의료 및 헬스케어 지수는 의료 및 헬스케어 업무에서 AI 모델의 성능을 측정하는 독립적인 벤치마크입니다. 임상 지식, 에이전트 지식 업무, 긴 환자 기록 추론, 환각 방지, 임상 추론, 에이전트 도구 사용을 평가합니다.
의료 및 헬스케어 지수는 의료 및 헬스케어 분야에서 모델의 성능을 평가하는 Artificial Analysis의 종합 벤치마크입니다. 의학, 공중 보건, 생의학 관련 전문 지식과 임상 진단 및 평가, 긴 환자 기록과 보험 청구 파일에 대한 추론, 환자 문서화, 약물 관리 등을 평가합니다.
의료 및 헬스케어 지수는 각 역량 하위 점수의 가중 평균으로 계산합니다. 하위 점수와 가중치는 의료 및 건강 지식 (30%), 에이전트형 지식 작업 (25%), 장문 맥락 추론 (15%), 환각 방지 (10%), 추론 (10%) 및 에이전트형 도구 사용 (10%)입니다.
의료 및 헬스케어 지수에는 AA-Omniscience 건강 정확도, GDPval-AA v2.1, AA-Briefcase v1.1, MLCR-AA, AA-Omniscience 건강 환각 방지, HLE 및 AutomationBench-AA 지원 및 운영이 포함됩니다.
결과가 공개된 모델 중 현재 Claude Opus 5.5 (Max, Default Fallback)의 의료 및 헬스케어 지수 점수가 61로 가장 높습니다. 모델 보기
의료 및 헬스케어 지수 점수가 높을수록 지수를 구성하는 벤치마크 전반에서 더 뛰어난 성능을 보인다는 뜻입니다. 특정 사용 사례에서는 종합 점수보다 개별 벤치마크 결과가 더 유용할 수 있습니다.