Artificial Analysis Search Index: benchmark y ranking de Search APIs

Las Search APIs permiten que los agentes de IA recuperen información de la web y fundamenten sus respuestas en fuentes actualizadas. Se utilizan en una amplia gama de aplicaciones agénticas, como la investigación profunda, la programación y el trabajo de conocimiento general, para mejorar la veracidad y la calidad de los resultados. Las Search APIs se diferencian en varios aspectos importantes: el índice web sobre el que buscan, la forma en que se presentan los resultados al modelo y los compromisos que asumen entre costo, velocidad y calidad de recuperación. Evaluamos 11 productos de Search API de 7 proveedores para ayudar a los desarrolladores a tomar decisiones mejor informadas al elegir una herramienta de búsqueda para su próximo agente de IA.

Para conocer las definiciones de los benchmarks, la puntuación y el tratamiento de los datos, consulta la página de metodología.

Datos más recientes: Aug 17, 2026.

Stirrup Stirrup agent
Benchmark tasks
Candidate model
GPT-5.6 Luna (medium)
turn 0 / 25
 
Search provider
 
web_search  
native payload, up to 10 results
Grading
Grader
avg F1 0.79
 
Search API providers
Constants
Same tasks, same model and grader (GPT-5.6 Luna, medium), same Stirrup harness (web_search · web_fetch · finish, max 25 turns). Only the search provider varies.

Aspectos destacados

Mejor resultado de Search API entre 3 benchmarks
Costo medio del modelo y de la Search API por tarea (USD) · Lower is better
Tiempo medio de modelo y de búsqueda por tarea · Lower is better

Resultados

Calidad de respuesta con búsqueda en benchmarks públicos.

Artificial Analysis Search Index

Media con igual ponderación del F1 de DeepSearchQA, la precisión en BrowseComp y la precisión en AA-Omniscience · Higher is better
Solo modelo (33)

The equal-weighted mean of each Search API provider's score on DeepSearchQA (average F1), BrowseComp (exact-answer accuracy), and AA-Omniscience (accuracy), shown on a 0-100 scale. Every result runs the same candidate answer model, harness, and settings. The only variable is the Search API provider.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

Costo

Costo del modelo candidato y de la Search API por tarea del benchmark y por consulta de búsqueda.

Costo por tarea

Costo medio del modelo candidato y de la Search API por tarea (USD) · Lower is better
Solo modelo ($0.0029)

Total cost of one benchmark task, split into the candidate answer model's tokens (input, cached, reasoning, and output) and the Search API provider's list price for the searches the model chose to run. The model-only baseline runs no searches, so its cost is token cost alone. Model cost can differ per provider due to differences in returned payloads for a given query, increased numbers of searches performed (due to not getting the right results) as well as increased reasoning token usage when search results are of lower quality.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

Artificial Analysis Search Index frente al costo por tarea

Artificial Analysis Search Index frente al costo total de modelo y búsqueda por tarea del benchmark (USD)
Most attractive quadrant
Pareto line

Total cost of one benchmark task, split into the candidate answer model's tokens (input, cached, reasoning, and output) and the Search API provider's list price for the searches the model chose to run. The model-only baseline runs no searches, so its cost is token cost alone. Model cost can differ per provider due to differences in returned payloads for a given query, increased numbers of searches performed (due to not getting the right results) as well as increased reasoning token usage when search results are of lower quality.

Latencia

Tiempo de modelo y tiempo de búsqueda por tarea del benchmark y por consulta de búsqueda.

Tiempo por tarea

Tiempo medio de modelo derivado y tiempo de búsqueda medido por tarea · Lower is better
Solo modelo (15.9s)

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

Artificial Analysis Search Index frente al tiempo por tarea

Artificial Analysis Search Index frente al tiempo medio de modelo y búsqueda por tarea del benchmark (segundos)
Most attractive quadrant
Pareto line

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

Detalles del ranking

Filas ordenables del benchmark público de Search APIs y desglose por benchmark.

Ranking público de Search APIs

Calidad, latencia y costo de Search API por proveedor, con desglose por benchmark.
12 filas
Proveedor
Parallel Search (advanced) logoParallel Search (advanced)
75
42
81
77
67
$47.93
$35.58
37.5s
Exa Search (auto) logoExa Search (auto)
74
41
78
74
70
$65.57
$61.58
27.8s
Firecrawl Search logoFirecrawl Search
73
40
74
74
73
$30.48
$44.94
56.9s
Parallel Search (basic) logoParallel Search (basic)
73
40
79
73
68
$45.14
$69.69
22.2s
Exa Search (fast) logoExa Search (fast)
68
35
76
61
69
$78.11
$80.53
22.6s
You.com Search logoYou.com Search
68
35
63
74
66
$68.93
$58.28
35.8s
Parallel Search (turbo) logoParallel Search (turbo)
67
34
70
75
56
$13.64
$46.57
20.7s
Keenable Search (pro) logoKeenable Search (pro)
67
34
70
66
65
$23.69
$73.20
24.8s
Keenable Search (realtime) logoKeenable Search (realtime)
67
34
70
64
66
$26.87
$63.67
16.8s
Tavily Search (basic) logoTavily Search (basic)
66
33
74
59
64
$126
$66.58
38.6s
Brave Search logoBrave Search
65
32
62
67
65
$71.61
$74.52
28.9s
Solo modelo
33
línea base
45
17
38
$2.92
15.9s

Search $/1k and Model $/1k are the average cost of 1,000 benchmark tasks broken down into search and model token costs. The model-only baseline runs no searches.

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

Tareas de ejemplo

Tareas de ejemplo representativas de cada benchmark que compone el Artificial Analysis Search Index.

DeepSearchQA

Broad research questions that need many searches. Answers are lists of items. An LLM grader scores each answer with an F1 score over the answer items. The full eval split has 900 tasks.

Representative example

I've completed the Desert Treasure quest on Old School Runescape, can you list all the names of the spells I can use to teleport into the wilderness that require more than four runes?

Answer: Dareeyak Teleport, Ghorrock Teleport

BrowseComp

Hard-to-find facts that need multi-hop browsing. We use a hard 200-sample subset from the full evaluation pool. The grader checks for an exact answer.

Representative example

I'm looking for the name and location of a structure which fulfills the following criteria: 1. Located in Eastern Australia. 2. Can be visited on foot. 3. Was re-built in 2016. 4. Is longer than 50m. 5. Can be seen from another similar structure. 6. Hosts a yearly dinner. 7. Was originally built for another use case, but has not served that use case for a number of years.

Answer: Shorncliffe Pier, Brisbane, Australia

AA-Omniscience

A private subset of 600 factual questions, balanced across 6 domains. The grader scores accuracy. The score shows how much search adds to the model's internal knowledge.

Representative example

In Karen Russell's short story "St. Lucy's Home for Girls Raised by Wolves," from which locality was the three-piece jazz band hired for the Debutante Ball?

Answer: West Toowoomba

The examples are representative problem types, not samples from the benchmarks.

Metodología e información adicional

Las métricas de Search API se agregan a partir de ejecuciones públicas de benchmarks e incluyen líneas base de solo modelo cuando están disponibles. Consulta la página de metodología para conocer los detalles de puntuación, latencia, costo y cobertura de benchmarks.

Preguntas frecuentes

Parallel Search (advanced) lidera el Artificial Analysis Search Index con 75 entre los 7 proveedores de Search API evaluados por Artificial Analysis.

Por tarea, Keenable Search (realtime) es la más rápida, con un tiempo medio por tarea de 16.8s (tiempo de modelo más tiempo de búsqueda). Por consulta de búsqueda individual, Keenable Search (realtime) es la más rápida, con un tiempo medio por consulta de búsqueda de 0.34s.

Parallel Search (turbo) tiene el menor costo de búsqueda medido, con $13.64 por cada 1000 tareas del benchmark.

La búsqueda aporta la mayor mejora de calidad medida en Parallel Search (advanced), elevando el Artificial Analysis Search Index en 42 frente al mismo modelo sin búsqueda (su línea base de solo modelo).

La calidad de respuesta de las Search APIs se agrega a partir de benchmarks públicos, entre ellos AA-Omniscience, BrowseComp y DeepSearchQA. Cada uno mide la capacidad de un modelo para encontrar, extraer o verificar información mediante búsqueda.

El mejor proveedor depende de tus prioridades. Usa los gráficos de calidad para comparar la precisión de las respuestas, los gráficos de costo para equilibrar calidad frente al costo de búsqueda y de modelo, y los gráficos de latencia para casos de uso en tiempo real. Los diagramas de dispersión de compromisos destacan a los proveedores situados en las fronteras de calidad-costo y calidad-latencia. Ver la metodología completa