Artificial Analysis Search Index:Search API 基准测试与排行榜

Search API 让 AI 智能体能够从网络获取信息,并让回答基于最新的来源。它们被广泛用于各类智能体应用,包括深度研究、编程和通用知识工作,以提升事实准确性和整体输出质量。Search API 在若干重要方面存在差异,包括其检索的底层网页索引、搜索结果呈现给模型的方式,以及在成本、速度和检索质量之间所做的取舍。我们对 7 家服务商的 11 款 Search API 产品进行了基准测试,帮助开发者在为下一个 AI 智能体挑选搜索工具时做出更明智的决策。

有关基准测试定义、评分方式和数据处理细节,请参阅方法论页面

最新数据:Aug 17, 2026。

Stirrup Stirrup agent
Benchmark tasks
Candidate model
GPT-5.6 Luna (medium)
turn 0 / 25
 
Search provider
 
web_search  
native payload, up to 10 results
Grading
Grader
avg F1 0.79
 
Search API providers
Constants
Same tasks, same model and grader (GPT-5.6 Luna, medium), same Stirrup harness (web_search · web_fetch · finish, max 25 turns). Only the search provider varies.

亮点

Search API 最佳结果(基准测试:3)
每任务的平均模型成本与 Search API 成本(美元) · Lower is better
每任务的平均模型耗时与搜索耗时 · Lower is better

结果

在各项公开基准测试中启用搜索后的回答质量。

Artificial Analysis Search Index

DeepSearchQA F1、BrowseComp 准确率与 AA-Omniscience 准确率的等权平均值 · Higher is better
仅模型(33)

The equal-weighted mean of each Search API provider's score on DeepSearchQA (average F1), BrowseComp (exact-answer accuracy), and AA-Omniscience (accuracy), shown on a 0-100 scale. Every result runs the same candidate answer model, harness, and settings. The only variable is the Search API provider.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

成本

每项基准测试任务和每次搜索查询的候选模型成本与 Search API 成本。

每任务成本

每任务的平均候选模型成本与 Search API 成本(美元) · Lower is better
仅模型($0.0029)

Total cost of one benchmark task, split into the candidate answer model's tokens (input, cached, reasoning, and output) and the Search API provider's list price for the searches the model chose to run. The model-only baseline runs no searches, so its cost is token cost alone. Model cost can differ per provider due to differences in returned payloads for a given query, increased numbers of searches performed (due to not getting the right results) as well as increased reasoning token usage when search results are of lower quality.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

Artificial Analysis Search Index 与每任务成本

Artificial Analysis Search Index 与每项基准测试任务的模型与搜索总成本(美元)
Most attractive quadrant
Pareto line

Total cost of one benchmark task, split into the candidate answer model's tokens (input, cached, reasoning, and output) and the Search API provider's list price for the searches the model chose to run. The model-only baseline runs no searches, so its cost is token cost alone. Model cost can differ per provider due to differences in returned payloads for a given query, increased numbers of searches performed (due to not getting the right results) as well as increased reasoning token usage when search results are of lower quality.

延迟

每项基准测试任务和每次搜索查询的模型耗时与搜索耗时。

每任务耗时

每任务的平均推算模型耗时与实测搜索耗时 · Lower is better
仅模型(15.9s)

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

Artificial Analysis Search Index 与每任务耗时

Artificial Analysis Search Index 与每项基准测试任务的平均模型与搜索耗时(秒)
Most attractive quadrant
Pareto line

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

排行榜详情

可排序的 Search API 公开基准测试数据行,以及各项基准测试的细分数据。

Search API 公开排行榜

按服务商划分的 Search API 质量、延迟、成本,以及各项基准测试的细分数据。
12 行
服务商
Parallel Search(advanced) logoParallel Search(advanced)
75
42
81
77
67
$47.93
$35.58
37.5s
Exa Search(auto) logoExa Search(auto)
74
41
78
74
70
$65.57
$61.58
27.8s
Firecrawl Search logoFirecrawl Search
73
40
74
74
73
$30.48
$44.94
56.9s
Parallel Search(basic) logoParallel Search(basic)
73
40
79
73
68
$45.14
$69.69
22.2s
Exa Search(fast) logoExa Search(fast)
68
35
76
61
69
$78.11
$80.53
22.6s
You.com Search logoYou.com Search
68
35
63
74
66
$68.93
$58.28
35.8s
Parallel Search(turbo) logoParallel Search(turbo)
67
34
70
75
56
$13.64
$46.57
20.7s
Keenable Search(pro) logoKeenable Search(pro)
67
34
70
66
65
$23.69
$73.20
24.8s
Keenable Search(realtime) logoKeenable Search(realtime)
67
34
70
64
66
$26.87
$63.67
16.8s
Tavily Search(basic) logoTavily Search(basic)
66
33
74
59
64
$126
$66.58
38.6s
Brave Search logoBrave Search
65
32
62
67
65
$71.61
$74.52
28.9s
仅模型
33
基线
45
17
38
$2.92
15.9s

Search $/1k and Model $/1k are the average cost of 1,000 benchmark tasks broken down into search and model token costs. The model-only baseline runs no searches.

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

示例任务

构成 Artificial Analysis Search Index 的各项基准测试的代表性示例任务。

DeepSearchQA

Broad research questions that need many searches. Answers are lists of items. An LLM grader scores each answer with an F1 score over the answer items. The full eval split has 900 tasks.

Representative example

I've completed the Desert Treasure quest on Old School Runescape, can you list all the names of the spells I can use to teleport into the wilderness that require more than four runes?

Answer: Dareeyak Teleport, Ghorrock Teleport

BrowseComp

Hard-to-find facts that need multi-hop browsing. We use a hard 200-sample subset from the full evaluation pool. The grader checks for an exact answer.

Representative example

I'm looking for the name and location of a structure which fulfills the following criteria: 1. Located in Eastern Australia. 2. Can be visited on foot. 3. Was re-built in 2016. 4. Is longer than 50m. 5. Can be seen from another similar structure. 6. Hosts a yearly dinner. 7. Was originally built for another use case, but has not served that use case for a number of years.

Answer: Shorncliffe Pier, Brisbane, Australia

AA-Omniscience

A private subset of 600 factual questions, balanced across 6 domains. The grader scores accuracy. The score shows how much search adds to the model's internal knowledge.

Representative example

In Karen Russell's short story "St. Lucy's Home for Girls Raised by Wolves," from which locality was the three-piece jazz band hired for the Debutante Ball?

Answer: West Toowoomba

The examples are representative problem types, not samples from the benchmarks.

方法论与更多信息

Search API 指标汇总自公开的基准测试运行结果,并在可用时包含仅模型基线。有关评分、延迟、成本和基准测试覆盖范围的细节,请参阅方法论页面

常见问题

在 Artificial Analysis 基准测试的 7 家 Search API 服务商中,Parallel Search(advanced) 以 75 位居 Artificial Analysis Search Index 首位。

按任务计算,Keenable Search(realtime) 最快,每任务平均耗时为 16.8s(模型耗时加搜索耗时)。按单次搜索查询计算,Keenable Search(realtime) 最快,每次搜索查询平均耗时为 0.34s。

Parallel Search(turbo) 的实测搜索成本最低,每 1,000 项基准测试任务为 $13.64。

搜索为 Parallel Search(advanced) 带来的实测质量提升最大,相比同一模型在不使用搜索时(其仅模型基线),Artificial Analysis Search Index 提高了 42。

Search API 的回答质量汇总自公开基准测试,包括 AA-Omniscience、BrowseComp和DeepSearchQA。每项基准测试都衡量模型借助搜索查找、提取或核实信息的能力。

最合适的服务商取决于你的优先事项。可用质量图表比较回答准确率,用成本图表在质量与搜索、模型成本之间取得平衡,用延迟图表评估实时场景。权衡散点图会标出位于质量-成本和质量-延迟前沿的服务商。 查看完整方法论