Artificial Analysis Search Index:Search API 基准测试与排行榜

Search API 让 AI 智能体能够从网络获取信息,并让回答基于最新的来源。它们被广泛用于各类智能体应用,包括深度研究、编程和通用知识工作,以提升事实准确性和整体输出质量。Search API 在若干重要方面存在差异,包括其检索的底层网页索引、搜索结果呈现给模型的方式,以及在成本、速度和检索质量之间所做的取舍。我们对 12 家服务商的 25 款 Search API 产品进行了基准测试,帮助开发者在为下一个 AI 智能体挑选搜索工具时做出更明智的决策。

有关基准测试定义、评分方式和数据处理细节,请参阅方法论页面。

最新数据:Sep 22, 2026。

亮点

Search API 最佳结果(基准测试:3)
每任务的平均模型成本与 Search API 成本(美元) · Lower is better
每任务的平均模型耗时与搜索耗时 · Lower is better

结果

在各项公开基准测试中启用搜索后的回答质量。

Artificial Analysis Search Index

DeepSearchQA F1、BrowseComp 准确率与 AA-Omniscience 准确率的等权平均值 · Higher is better
仅模型(33)

成本

每项基准测试任务和每次搜索查询的候选模型成本与 Search API 成本。

每任务成本

每任务的平均候选模型成本与 Search API 成本(美元) · Lower is better
仅模型($0.0029)

Artificial Analysis Search Index 与每任务成本

Artificial Analysis Search Index 与每项基准测试任务的模型与搜索总成本(美元)
Most attractive quadrant
Pareto line

延迟

每项基准测试任务和每次搜索查询的模型耗时与搜索耗时。

每任务耗时

每任务的平均推算模型耗时与实测搜索耗时 · Lower is better
仅模型(20.3s)

Artificial Analysis Search Index 与每任务耗时

Artificial Analysis Search Index 与每项基准测试任务的平均模型与搜索耗时(秒)
Most attractive quadrant
Pareto line

Token

每项基准测试任务的候选模型输入与输出 token 数,以及质量与 token 总用量的关系。

每任务 token 总数

每项基准测试任务中候选模型的平均输入、推理与回答 token 数 · Lower is better
仅模型(2.6k)

Artificial Analysis Search Index 与每任务 token 总数

Artificial Analysis Search Index 与每项基准测试任务的候选模型 token 总数
Most attractive quadrant
Pareto line

搜索

候选模型在每项基准测试任务中执行的搜索查询次数。

每任务搜索查询次数

候选模型在每项基准测试任务中执行的平均搜索查询次数

排行榜详情

可排序的 Search API 公开基准测试数据行,以及各项基准测试的细分数据。

Search API 公开排行榜

按服务商划分的 Search API 质量、延迟、成本,以及各项基准测试的细分数据。
13 行
服务商
Perplexity Search(medium) logoPerplexity Search(medium)
80
47
81
87
72
$62.30
$8.04
27.1s
Perplexity Search(high) logoPerplexity Search(high)
79
46
81
86
71
$56.93
$9.72
27.8s
Octen Search(highlights) logoOcten Search(highlights)
77
44
80
86
66
$9.07
$14.92
15.4s
Parallel Search(advanced) logoParallel Search(advanced)
75
42
81
77
67
$47.93
$11.29
40.9s
Brave Search(LLM context) logoBrave Search(LLM context)
75
42
78
77
69
$61.96
$17.46
22.2s
You.com Search(highlights) logoYou.com Search(highlights)
74
41
77
77
69
$47.54
$19.58
19.1s
Nimble Search(standard) logoNimble Search(standard)
74
41
76
75
71
$46.14
$19.57
41.0s
Exa Search(auto) logoExa Search(auto)
74
41
78
74
70
$65.57
$17.73
31.2s
Firecrawl Search logoFirecrawl Search
73
40
74
74
73
$30.48
$12.09
60.9s
Parallel Search(basic) logoParallel Search(basic)
73
40
79
73
68
$45.14
$21.65
25.8s
TinyFish Search(web) logoTinyFish Search(web)
71
38
67
75
71
$0
$9.30
57.5s
Exa Search(instant) logoExa Search(instant)
68
35
73
65
67
$72.23
$21.44
18.4s
仅模型
33
基线
45
17
38
—
$2.92
20.3s

示例任务

构成 Artificial Analysis Search Index 的各项基准测试的代表性示例任务。

DeepSearchQA

Broad research questions that need many searches. Answers are lists of items. An LLM grader scores each answer with an F1 score over the answer items. The full eval split has 900 tasks.

Representative example

I've completed the Desert Treasure quest on Old School Runescape, can you list all the names of the spells I can use to teleport into the wilderness that require more than four runes?

Answer: Dareeyak Teleport, Ghorrock Teleport

BrowseComp

Hard-to-find facts that need multi-hop browsing. We use a hard 200-sample subset from the full evaluation pool. The grader checks for an exact answer.

Representative example

I'm looking for the name and location of a structure which fulfills the following criteria: 1. Located in Eastern Australia. 2. Can be visited on foot. 3. Was re-built in 2016. 4. Is longer than 50m. 5. Can be seen from another similar structure. 6. Hosts a yearly dinner. 7. Was originally built for another use case, but has not served that use case for a number of years.

Answer: Shorncliffe Pier, Brisbane, Australia

AA-Omniscience

A private subset of 600 factual questions, balanced across 6 domains. The grader scores accuracy. The score shows how much search adds to the model's internal knowledge.

Representative example

In Karen Russell's short story "St. Lucy's Home for Girls Raised by Wolves," from which locality was the three-piece jazz band hired for the Debutante Ball?

Answer: West Toowoomba

The examples are representative problem types, not samples from the benchmarks.

方法论与更多信息

Search API 指标汇总自公开的基准测试运行结果,并在可用时包含仅模型基线。有关评分、延迟、成本和基准测试覆盖范围的细节,请参阅方法论页面。

常见问题