Todos los índices de capacidades

Healthcare & Medical Index

Assesses model performance across the healthcare and medical domain. Capabilities evaluated include domain-specific knowledge (medicine, public health, biomedical sciences), clinical diagnosis and assessment, reasoning over long patient records and claims files, patient documentation, medication management, and more.

Ver flujos de trabajo representativos

The Artificial Analysis Healthcare & Medical Index combines performance across benchmarks chosen for clinical and healthcare-support work, spanning medical knowledge, clinical reasoning, long-context reasoning over patient records, agentic workflows, and non-hallucination. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.

This composite metric provides a single score for tracking model performance across healthcare tasks. Artificial Analysis ejecuta de forma independiente todas las evaluaciones subyacentes. Consulta nuestra metodología de Benchmarking de Inteligencia para saber cómo se realizan las evaluaciones.

CapacidadPesoEvaluaciones
Medical & Health Knowledge30 %AA-Omniscience Health Accuracy
Agentic Knowledge Work25 %GDPval-AA v2
Long-Context Reasoning15 %MLCR-AA
Non-Hallucination10 %AA-Omniscience Non-Hallucination
Reasoning10 %HLE
Agentic Customer Interaction10 %𝜏³-Banking

Puntuación

Artificial Analysis Healthcare & Medical Index

Incorporates 5 evaluations: AA-Omniscience, GDPval-AA v2, MLCR-AA, Humanity's Last Exam, 𝜏³-Banking · Higher is better

Artificial Analysis Healthcare & Medical Index: desglose de capacidades

Incorporates 5 evaluations: AA-Omniscience, GDPval-AA v2, MLCR-AA, Humanity's Last Exam, 𝜏³-Banking · Desglosado por contribución

Desglose de capacidades

Artificial Analysis Healthcare & Medical Index: Medical & Health Knowledge

Incorporates 1 evaluation: AA-Omniscience · Higher is better

Flujos de trabajo representativos

Flujos de trabajo reales que ponen a prueba las capacidades a las que Healthcare & Medical Index da mayor peso.

Patient diagnosis & clinical assessmentMedical & Health KnowledgeLong-Context ReasoningNon-HallucinationReasoning

Ejemplo: Reassess a returning patient with worsening symptoms against the original EHR workup to build a differential from the new labs and imaging and surface alternative diagnoses the findings point to.

Treatment administration & proceduresMedical & Health KnowledgeNon-Hallucination

Ejemplo: A surgical team that encounters unexpected anatomy mid-laparoscopic-procedure. Retrieve comparable case reports and imaging precedents and quickly output findings relevant to their immediate decision.

Patient documentation & chartingMedical & Health KnowledgeAgentic Knowledge Work

Ejemplo: Turn a clinician's dictated notes from a follow-up visit into a structured SOAP note, pulling the patient's active problems and relevant history from the existing chart, placing each finding in the right section, and flagging the gaps the next provider would need filled.

Medication management & pharmacy coordinationMedical & Health KnowledgeNon-HallucinationReasoning

Ejemplo: Calculate a child's per-dose amount from their measurements and the prescriber's notes against the available suspension concentration, convert it to the millilitres to measure at each dose, and produce caregiver instructions that keep the total within the safe daily range.

Continuing education & researchMedical & Health KnowledgeAgentic Knowledge WorkLong-Context ReasoningReasoning

Ejemplo: Evaluate whether a dermatology team should adopt a newer procedure backed by emerging but limited long-term evidence to summarise the published trials and safety data, compare outcomes against the current standard of care, and outline the open questions the team still needs to resolve.

Medical records review & claimsLong-Context ReasoningNon-HallucinationReasoning

Ejemplo: Work a several-hundred-page medical record assembled from multiple providers to reconstruct the treatment timeline, identify which encounters relate to the injury in question, and answer reviewer questions with citations to the underlying documents.

Patient education & care coordinationMedical & Health KnowledgeAgentic Knowledge WorkAgentic Customer Interaction

Ejemplo: Turn a patient's after-visit summary into plain-language, step-by-step home-care instructions in their preferred language, anticipate the questions they are most likely to ask, and confirm the follow-up appointment and how to reach the clinic with concerns.

Fecha de lanzamiento

Artificial Analysis Healthcare & Medical Index vs. fecha de lanzamiento

Most attractive region

Costo

Artificial Analysis Healthcare & Medical Index: costo por tarea

Costo medio por tarea (USD), desglosado en tokens de entrada, aciertos y escrituras de caché, razonamiento y respuesta

Average cost per task in the index. Costs are split by input, cache hit, cache write, reasoning, and answer token pricing where canonical token counts are available.

Artificial Analysis Healthcare & Medical Index: costo total

Costo total (USD) de ejecutar el índice

The cost to run the index, calculated using the model's input and output token pricing and the number of tokens used.

Velocidad

Artificial Analysis Healthcare & Medical Index: tiempo por tarea

Weighted average decode time (minutes) per task; excludes TTFT and overhead time · Lower is better

The weighted average time (minutes) per index task. This is calculated by dividing output tokens per task by output speed, weighted by the relative weights of each benchmark in the index.

Tokens de salida

Artificial Analysis Healthcare & Medical Index: tokens de salida por tarea

Tokens de salida utilizados para ejecutar una tarea, desglosados en tokens de razonamiento y respuesta

The average number of answer and reasoning tokens produced per benchmark task in this index.

Preguntas frecuentes

Según el Healthcare & Medical Index de Artificial Analysis, los modelos de IA con mejor desempeño en trabajos de salud y medicina son actualmente Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (56), Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (52) y Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (52). La clasificación se actualiza a medida que se lanzan nuevos modelos.

Sí. El Healthcare & Medical Index de Artificial Analysis es un benchmark independiente que mide el desempeño de los modelos de IA en trabajos de salud y medicina. Evalúa conocimientos y razonamiento clínicos, razonamiento con historiales extensos de pacientes, flujos de trabajo agénticos y ausencia de alucinaciones.

El Healthcare & Medical Index es un benchmark compuesto de Artificial Analysis que evalúa el desempeño de los modelos en salud y medicina. Las capacidades evaluadas incluyen conocimientos específicos (medicina, salud pública y ciencias biomédicas), diagnóstico y evaluación clínica, razonamiento sobre historiales extensos y expedientes de reclamaciones, documentación de pacientes, gestión de medicamentos y más.

El Healthcare & Medical Index se calcula como el promedio ponderado de las puntuaciones de sus capacidades. Estas son las puntuaciones y sus pesos: Medical & Health Knowledge (30 %), Agentic Knowledge Work (25 %), Long-Context Reasoning (15 %), Non-Hallucination (10 %), Reasoning (10 %) y Agentic Customer Interaction (10 %).

El Healthcare & Medical Index incluye AA-Omniscience Health Accuracy, GDPval-AA v2, MLCR-AA, AA-Omniscience Non-Hallucination, HLE y 𝜏³-Banking.

Actualmente, Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) tiene la puntuación más alta en el Healthcare & Medical Index: 56 entre los modelos con resultados publicados. Ver modelo

Una puntuación más alta en el Healthcare & Medical Index indica un mejor desempeño general en los benchmarks que componen el índice. Para un caso de uso específico, los resultados de cada benchmark pueden ser más informativos que la puntuación compuesta.