Healthcare & Medical Index
Assesses model performance across the healthcare and medical domain. Capabilities evaluated include domain-specific knowledge (medicine, public health, biomedical sciences), clinical diagnosis and assessment, reasoning over long patient records and claims files, patient documentation, medication management, and more.
Ver fluxos de trabalho representativosThe Artificial Analysis Healthcare & Medical Index combines performance across benchmarks chosen for clinical and healthcare-support work, spanning medical knowledge, clinical reasoning, long-context reasoning over patient records, agentic workflows, and non-hallucination. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across healthcare tasks. Todos os benchmarks subjacentes são executados de forma independente pela Artificial Analysis. Consulte nossa metodologia de benchmarking de inteligência para saber como as avaliações são realizadas.
| Capacidade | Peso | Avaliações |
|---|---|---|
| Medical & Health Knowledge | 30% | AA-Omniscience Health Accuracy |
| Agentic Knowledge Work | 25% | GDPval-AA v2 e AA-Briefcase |
| Long-Context Reasoning | 15% | MLCR-AA |
| Non-Hallucination | 10% | AA-Omniscience Health Non-Hallucination |
| Reasoning | 10% | HLE |
| Agentic Tool Use | 10% | AutomationBench-AA Support & Operations |
Pontuação
Artificial Analysis Healthcare & Medical Index
Artificial Analysis Healthcare & Medical Index: detalhamento das capacidades
Detalhamento das capacidades
Artificial Analysis Healthcare & Medical Index: Medical & Health Knowledge
Fluxos de trabalho representativos
Fluxos de trabalho reais que exercitam as capacidades às quais o Healthcare & Medical Index atribui mais peso.
Exemplo: Reassess a returning patient with worsening symptoms against the original EHR workup to build a differential from the new labs and imaging and surface alternative diagnoses the findings point to.
Exemplo: A surgical team that encounters unexpected anatomy mid-laparoscopic-procedure. Retrieve comparable case reports and imaging precedents and quickly output findings relevant to their immediate decision.
Exemplo: Turn a clinician's dictated notes from a follow-up visit into a structured SOAP note, pulling the patient's active problems and relevant history from the existing chart, placing each finding in the right section, and flagging the gaps the next provider would need filled.
Exemplo: Calculate a child's per-dose amount from their measurements and the prescriber's notes against the available suspension concentration, convert it to the millilitres to measure at each dose, and produce caregiver instructions that keep the total within the safe daily range.
Exemplo: Evaluate whether a dermatology team should adopt a newer procedure backed by emerging but limited long-term evidence to summarise the published trials and safety data, compare outcomes against the current standard of care, and outline the open questions the team still needs to resolve.
Exemplo: Work a several-hundred-page medical record assembled from multiple providers to reconstruct the treatment timeline, identify which encounters relate to the injury in question, and answer reviewer questions with citations to the underlying documents.
Exemplo: Turn a patient's after-visit summary into plain-language, step-by-step home-care instructions in their preferred language, anticipate the questions they are most likely to ask, and confirm the follow-up appointment and how to reach the clinic with concerns.
Custo
Artificial Analysis Healthcare & Medical Index: custo por tarefa
Artificial Analysis Healthcare & Medical Index vs. custo por tarefa
Velocidade
Artificial Analysis Healthcare & Medical Index: tempo por tarefa
Tokens de saída
Artificial Analysis Healthcare & Medical Index: tokens de saída por tarefa
Data de lançamento
Artificial Analysis Healthcare & Medical Index vs. data de lançamento
Perguntas frequentes
Segundo o Healthcare & Medical Index da Artificial Analysis, os modelos de IA com melhor desempenho em trabalhos de saúde e medicina atualmente são Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (58), Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (54) e Claude Opus 5 (Adaptive Reasoning, Max Effort) (53). O ranking é atualizado à medida que novos modelos são lançados.
Sim. O Healthcare & Medical Index da Artificial Analysis é um benchmark independente do desempenho de modelos de IA em trabalhos de saúde e medicina. Ele mede conhecimento clínico, trabalho de conhecimento agêntico, raciocínio sobre prontuários longos, ausência de alucinações, raciocínio clínico e uso agêntico de ferramentas.
O Healthcare & Medical Index é um benchmark composto da Artificial Analysis que avalia o desempenho dos modelos em saúde e medicina. As capacidades avaliadas incluem conhecimento específico (medicina, saúde pública e ciências biomédicas), diagnóstico e avaliação clínica, raciocínio sobre prontuários extensos e arquivos de sinistros, documentação de pacientes, gestão de medicamentos e muito mais.
O Healthcare & Medical Index é calculado como a média ponderada das pontuações de suas capacidades. Estas são as pontuações e seus pesos: Medical & Health Knowledge (30%), Agentic Knowledge Work (25%), Long-Context Reasoning (15%), Non-Hallucination (10%), Reasoning (10%) e Agentic Tool Use (10%).
O Healthcare & Medical Index inclui AA-Omniscience Health Accuracy, GDPval-AA v2, AA-Briefcase, MLCR-AA, AA-Omniscience Health Non-Hallucination, HLE e AutomationBench-AA Support & Operations.
Atualmente, Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) tem a maior pontuação no Healthcare & Medical Index: 58 entre os modelos com resultados publicados. Ver modelo
Uma pontuação mais alta no Healthcare & Medical Index indica melhor desempenho geral nos benchmarks que compõem o índice. Para um caso de uso específico, os resultados de cada benchmark podem ser mais informativos que a pontuação composta.