Artificial Analysis Search Index: Search API のベンチマークとリーダーボード

Search API は、AI エージェントがウェブから情報を取得し、最新の情報源に基づいて回答を組み立てることを可能にします。ディープリサーチ、コーディング、一般的なナレッジワークまで幅広いエージェント型アプリケーションで利用され、事実性と出力全体の品質向上に役立っています。Search API は、検索対象となるウェブインデックス、検索結果をモデルに提示する方法、そしてコスト・速度・検索品質のトレードオフなど、いくつかの重要な点で異なります。Artificial Analysis では 10 社のプロバイダーが提供する 20 件の Search API 製品をベンチマークし、開発者が次の AI エージェント向けに検索ツールを選ぶ際、より的確な判断を下せるよう支援します。

ベンチマークの定義、採点方法、データの取り扱いの詳細は方法論ページをご覧ください。

最新データ: Sep 8, 2026。

ハイライト

Search API のトップ結果(対象ベンチマーク: 3)
タスクあたりのモデル費用と Search API 費用の平均(USD) · Lower is better
タスクあたりのモデル時間と検索時間の平均 · Lower is better

結果

公開ベンチマーク全体における、検索を有効にした場合の回答品質。

Artificial Analysis Search Index

DeepSearchQA F1、BrowseComp 精度、AA-Omniscience 精度を等しい重みで平均した値 · Higher is better
モデルのみ(33)

費用

ベンチマークタスクあたり、および検索クエリあたりの候補モデル費用と Search API 費用。

タスクあたりの費用

タスクあたりの候補モデル費用と Search API 費用の平均(USD) · Lower is better
モデルのみ($0.0029)

Artificial Analysis Search Index とタスクあたりの費用

Artificial Analysis Search Index と、ベンチマークタスクあたりのモデル費用と検索費用の合計(USD)
Most attractive quadrant
Pareto line

レイテンシ

ベンチマークタスクあたり、および検索クエリあたりのモデル時間と検索時間。

タスクあたりの時間

タスクあたりの推定モデル時間と実測検索時間の平均 · Lower is better
モデルのみ(21.3s)

Artificial Analysis Search Index とタスクあたりの時間

Artificial Analysis Search Index と、ベンチマークタスクあたりの平均モデル時間と検索時間(秒)
Most attractive quadrant
Pareto line

トークン

ベンチマークタスクあたりの候補モデルの入力・出力トークン数と、トークン使用量全体に対する品質。

タスクあたりの合計トークン数

ベンチマークタスクあたりの候補モデルの入力・推論・回答トークンの平均 · Lower is better
モデルのみ(2.6k)

Artificial Analysis Search Index とタスクあたりの合計トークン数

Artificial Analysis Search Index と、ベンチマークタスクあたりの候補モデルの合計トークン数
Most attractive quadrant
Pareto line

検索クエリ

ベンチマークタスクあたりに候補モデルが実行する検索クエリ数。

タスクあたりの検索クエリ数

ベンチマークタスクあたりに候補モデルが実行する検索クエリの平均回数

リーダーボードの詳細

並べ替え可能な Search API 公開ベンチマークの行と、ベンチマークごとの内訳。

Search API 公開リーダーボード

プロバイダー単位の Search API 品質・レイテンシ・費用と、ベンチマークごとの内訳。
13 行
プロバイダー
Perplexity Search(medium) logoPerplexity Search(medium)
80
47
81
87
72
$62.30
$29.09
27.7s
Perplexity Search(high) logoPerplexity Search(high)
79
46
81
86
71
$56.93
$34.44
28.5s
Octen Search(highlights) logoOcten Search(highlights)
77
44
80
86
66
$9.07
$49.15
16.1s
Perplexity Search(low) logoPerplexity Search(low)
77
44
76
85
70
$77.20
$27.58
35.8s
Parallel Search(advanced) logoParallel Search(advanced)
75
42
81
77
67
$47.93
$35.58
41.6s
Brave Search(LLM context) logoBrave Search(LLM context)
75
42
78
77
69
$61.96
$67.57
22.9s
You.com Search(highlights) logoYou.com Search(highlights)
74
41
77
77
69
$47.54
$69.94
19.7s
Exa Search(auto) logoExa Search(auto)
74
41
78
74
70
$65.57
$61.58
31.9s
Firecrawl Search logoFirecrawl Search
73
40
74
74
73
$30.48
$44.94
61.7s
Parallel Search(basic) logoParallel Search(basic)
73
40
79
73
68
$45.14
$69.69
26.5s
Parallel Search(fast) logoParallel Search(fast)
73
40
80
72
68
$8.41
$59.67
18.5s
TinyFish Search(web) logoTinyFish Search(web)
71
38
67
75
71
$0
$34.55
58.5s
モデルのみ
33
ベースライン
45
17
38
$2.92
21.3s

タスクの例

Artificial Analysis Search Index を構成する各ベンチマークの代表的なタスク例。

DeepSearchQA

Broad research questions that need many searches. Answers are lists of items. An LLM grader scores each answer with an F1 score over the answer items. The full eval split has 900 tasks.

Representative example

I've completed the Desert Treasure quest on Old School Runescape, can you list all the names of the spells I can use to teleport into the wilderness that require more than four runes?

Answer: Dareeyak Teleport, Ghorrock Teleport

BrowseComp

Hard-to-find facts that need multi-hop browsing. We use a hard 200-sample subset from the full evaluation pool. The grader checks for an exact answer.

Representative example

I'm looking for the name and location of a structure which fulfills the following criteria: 1. Located in Eastern Australia. 2. Can be visited on foot. 3. Was re-built in 2016. 4. Is longer than 50m. 5. Can be seen from another similar structure. 6. Hosts a yearly dinner. 7. Was originally built for another use case, but has not served that use case for a number of years.

Answer: Shorncliffe Pier, Brisbane, Australia

AA-Omniscience

A private subset of 600 factual questions, balanced across 6 domains. The grader scores accuracy. The score shows how much search adds to the model's internal knowledge.

Representative example

In Karen Russell's short story "St. Lucy's Home for Girls Raised by Wolves," from which locality was the three-piece jazz band hired for the Debutante Ball?

Answer: West Toowoomba

The examples are representative problem types, not samples from the benchmarks.

方法論と詳細情報

Search API の指標は公開ベンチマークの実行結果を集計したもので、利用できる場合はモデルのみのベースラインも含みます。採点、レイテンシ、費用、ベンチマークの網羅範囲の詳細は方法論ページをご覧ください。

よくある質問