Artificial Analysis Search Index: Search API Benchmark & Leaderboard

Search APIs allow AI agents to retrieve information from the web, grounding their responses in up-to-date sources. They are used across a wide range of agentic applications, including deep research, coding, and general knowledge work, to improve factuality and overall outputs. Search APIs differ in several important ways, including the underlying web index they search, how search results are presented to the model, and the tradeoffs they make between cost, speed, and retrieval quality. We benchmark 20 Search API products across 10 providers to help developers make more informed decisions when selecting a search tool for their next AI agent.

For benchmark definitions, scoring, and data handling details, see the methodology page.

Latest data: Sep 8, 2026.

Highlights

Top Search API result across 3 benchmarks
Average model and Search API cost per task (USD) · Lower is better
Average model and search time per task · Lower is better

Results

Search-enabled answer quality across public benchmarks.

Artificial Analysis Search Index

Equal-weighted mean of DeepSearchQA F1, BrowseComp accuracy, and AA-Omniscience accuracy · Higher is better
Model only (33)

Cost

Candidate model cost and Search API cost per benchmark task and per search query.

Cost per Task

Average candidate model cost and Search API cost per task (USD) · Lower is better
Model only ($0.0029)

Artificial Analysis Search Index vs. Cost per Task

Artificial Analysis Search Index vs. total model and search cost per benchmark task (USD)
Most attractive quadrant
Pareto line

Latency

Model time and search time per benchmark task and per search query.

Time per Task

Average derived model time and measured search time per task · Lower is better
Model only (23.8s)

Artificial Analysis Search Index vs. Time per Task

Artificial Analysis Search Index vs. average model and search time per benchmark task (seconds)
Most attractive quadrant
Pareto line

Tokens

Candidate model input and output tokens per benchmark task, and quality against total token usage.

Total Tokens per Task

Average candidate model input, reasoning, and answer tokens per benchmark task · Lower is better
Model only (2.6k)

Artificial Analysis Search Index vs. Total Tokens per Task

Artificial Analysis Search Index vs. total candidate model tokens per benchmark task
Most attractive quadrant
Pareto line

Searches

Search queries the candidate model runs per benchmark task.

Search Queries per Task

Average number of search queries the candidate model makes per benchmark task

Leaderboard Details

Sortable public Search API benchmark rows and per benchmark breakdown.

Search API Public Leaderboard

Provider-level Search API quality, latency, cost, and per benchmark breakdown.
13 rows
Provider
Perplexity Search (medium) logoPerplexity Search (medium)
80
47
81
87
72
$62.30
$29.09
29.4s
Perplexity Search (high) logoPerplexity Search (high)
79
46
81
86
71
$56.93
$34.44
30.3s
Octen Search (highlights) logoOcten Search (highlights)
77
44
80
86
66
$9.07
$49.15
17.7s
Perplexity Search (low) logoPerplexity Search (low)
77
44
76
85
70
$77.20
$27.58
38.1s
Parallel Search (advanced) logoParallel Search (advanced)
75
42
81
77
67
$47.93
$35.58
43.5s
Brave Search (LLM context) logoBrave Search (LLM context)
75
42
78
77
69
$61.96
$67.57
25.0s
You.com Search (highlights) logoYou.com Search (highlights)
74
41
77
77
69
$47.54
$69.94
21.2s
Exa Search (auto) logoExa Search (auto)
74
41
78
74
70
$65.57
$61.58
33.9s
Firecrawl Search logoFirecrawl Search
73
40
74
74
73
$30.48
$44.94
63.9s
Parallel Search (basic) logoParallel Search (basic)
73
40
79
73
68
$45.14
$69.69
28.6s
Parallel Search (fast) logoParallel Search (fast)
73
40
80
72
68
$8.41
$59.67
19.8s
TinyFish Search (web) logoTinyFish Search (web)
71
38
67
75
71
$0
$34.55
61.3s
Model only
33
baseline
45
17
38
$2.92
23.8s

Example Tasks

Representative example tasks for each benchmark behind the Artificial Analysis Search Index.

DeepSearchQA

Broad research questions that need many searches. Answers are lists of items. An LLM grader scores each answer with an F1 score over the answer items. The full eval split has 900 tasks.

Representative example

I've completed the Desert Treasure quest on Old School Runescape, can you list all the names of the spells I can use to teleport into the wilderness that require more than four runes?

Answer: Dareeyak Teleport, Ghorrock Teleport

BrowseComp

Hard-to-find facts that need multi-hop browsing. We use a hard 200-sample subset from the full evaluation pool. The grader checks for an exact answer.

Representative example

I'm looking for the name and location of a structure which fulfills the following criteria: 1. Located in Eastern Australia. 2. Can be visited on foot. 3. Was re-built in 2016. 4. Is longer than 50m. 5. Can be seen from another similar structure. 6. Hosts a yearly dinner. 7. Was originally built for another use case, but has not served that use case for a number of years.

Answer: Shorncliffe Pier, Brisbane, Australia

AA-Omniscience

A private subset of 600 factual questions, balanced across 6 domains. The grader scores accuracy. The score shows how much search adds to the model's internal knowledge.

Representative example

In Karen Russell's short story "St. Lucy's Home for Girls Raised by Wolves," from which locality was the three-piece jazz band hired for the Debutante Ball?

Answer: West Toowoomba

The examples are representative problem types, not samples from the benchmarks.

Methodology & Further Information

Search API metrics are aggregated from public benchmark runs and include model-only baselines where available. See the methodology page for scoring, latency, cost, and benchmark coverage details.

Frequently Asked Questions