Artificial Analysis Search Index: Search API のベンチマークとリーダーボード

Search API は、AI エージェントがウェブから情報を取得し、最新の情報源に基づいて回答を組み立てることを可能にします。ディープリサーチ、コーディング、一般的なナレッジワークまで幅広いエージェント型アプリケーションで利用され、事実性と出力全体の品質向上に役立っています。Search API は、検索対象となるウェブインデックス、検索結果をモデルに提示する方法、そしてコスト・速度・検索品質のトレードオフなど、いくつかの重要な点で異なります。Artificial Analysis では 7 社のプロバイダーが提供する 11 件の Search API 製品をベンチマークし、開発者が次の AI エージェント向けに検索ツールを選ぶ際、より的確な判断を下せるよう支援します。

ベンチマークの定義、採点方法、データの取り扱いの詳細は方法論ページをご覧ください。

最新データ: Aug 17, 2026。

Stirrup Stirrup agent
Benchmark tasks
Candidate model
GPT-5.6 Luna (medium)
turn 0 / 25
 
Search provider
 
web_search  
native payload, up to 10 results
Grading
Grader
avg F1 0.79
 
Search API providers
Constants
Same tasks, same model and grader (GPT-5.6 Luna, medium), same Stirrup harness (web_search · web_fetch · finish, max 25 turns). Only the search provider varies.

ハイライト

Search API のトップ結果(対象ベンチマーク: 3)
タスクあたりのモデル費用と Search API 費用の平均(USD) · Lower is better
タスクあたりのモデル時間と検索時間の平均 · Lower is better

結果

公開ベンチマーク全体における、検索を有効にした場合の回答品質。

Artificial Analysis Search Index

DeepSearchQA F1、BrowseComp 精度、AA-Omniscience 精度を等しい重みで平均した値 · Higher is better
モデルのみ(33)

The equal-weighted mean of each Search API provider's score on DeepSearchQA (average F1), BrowseComp (exact-answer accuracy), and AA-Omniscience (accuracy), shown on a 0-100 scale. Every result runs the same candidate answer model, harness, and settings. The only variable is the Search API provider.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

費用

ベンチマークタスクあたり、および検索クエリあたりの候補モデル費用と Search API 費用。

タスクあたりの費用

タスクあたりの候補モデル費用と Search API 費用の平均(USD) · Lower is better
モデルのみ($0.0029)

Total cost of one benchmark task, split into the candidate answer model's tokens (input, cached, reasoning, and output) and the Search API provider's list price for the searches the model chose to run. The model-only baseline runs no searches, so its cost is token cost alone. Model cost can differ per provider due to differences in returned payloads for a given query, increased numbers of searches performed (due to not getting the right results) as well as increased reasoning token usage when search results are of lower quality.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

Artificial Analysis Search Index とタスクあたりの費用

Artificial Analysis Search Index と、ベンチマークタスクあたりのモデル費用と検索費用の合計(USD)
Most attractive quadrant
Pareto line

Total cost of one benchmark task, split into the candidate answer model's tokens (input, cached, reasoning, and output) and the Search API provider's list price for the searches the model chose to run. The model-only baseline runs no searches, so its cost is token cost alone. Model cost can differ per provider due to differences in returned payloads for a given query, increased numbers of searches performed (due to not getting the right results) as well as increased reasoning token usage when search results are of lower quality.

レイテンシ

ベンチマークタスクあたり、および検索クエリあたりのモデル時間と検索時間。

タスクあたりの時間

タスクあたりの推定モデル時間と実測検索時間の平均 · Lower is better
モデルのみ(15.9s)

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

Artificial Analysis Search Index とタスクあたりの時間

Artificial Analysis Search Index と、ベンチマークタスクあたりの平均モデル時間と検索時間(秒)
Most attractive quadrant
Pareto line

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

リーダーボードの詳細

並べ替え可能な Search API 公開ベンチマークの行と、ベンチマークごとの内訳。

Search API 公開リーダーボード

プロバイダー単位の Search API 品質・レイテンシ・費用と、ベンチマークごとの内訳。
12 行
プロバイダー
Parallel Search(advanced) logoParallel Search(advanced)
75
42
81
77
67
$47.93
$35.58
37.5s
Exa Search(auto) logoExa Search(auto)
74
41
78
74
70
$65.57
$61.58
27.8s
Firecrawl Search logoFirecrawl Search
73
40
74
74
73
$30.48
$44.94
56.9s
Parallel Search(basic) logoParallel Search(basic)
73
40
79
73
68
$45.14
$69.69
22.2s
Exa Search(fast) logoExa Search(fast)
68
35
76
61
69
$78.11
$80.53
22.6s
You.com Search logoYou.com Search
68
35
63
74
66
$68.93
$58.28
35.8s
Parallel Search(turbo) logoParallel Search(turbo)
67
34
70
75
56
$13.64
$46.57
20.7s
Keenable Search(pro) logoKeenable Search(pro)
67
34
70
66
65
$23.69
$73.20
24.8s
Keenable Search(realtime) logoKeenable Search(realtime)
67
34
70
64
66
$26.87
$63.67
16.8s
Tavily Search(basic) logoTavily Search(basic)
66
33
74
59
64
$126
$66.58
38.6s
Brave Search logoBrave Search
65
32
62
67
65
$71.61
$74.52
28.9s
モデルのみ
33
ベースライン
45
17
38
$2.92
15.9s

Search $/1k and Model $/1k are the average cost of 1,000 benchmark tasks broken down into search and model token costs. The model-only baseline runs no searches.

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

タスクの例

Artificial Analysis Search Index を構成する各ベンチマークの代表的なタスク例。

DeepSearchQA

Broad research questions that need many searches. Answers are lists of items. An LLM grader scores each answer with an F1 score over the answer items. The full eval split has 900 tasks.

Representative example

I've completed the Desert Treasure quest on Old School Runescape, can you list all the names of the spells I can use to teleport into the wilderness that require more than four runes?

Answer: Dareeyak Teleport, Ghorrock Teleport

BrowseComp

Hard-to-find facts that need multi-hop browsing. We use a hard 200-sample subset from the full evaluation pool. The grader checks for an exact answer.

Representative example

I'm looking for the name and location of a structure which fulfills the following criteria: 1. Located in Eastern Australia. 2. Can be visited on foot. 3. Was re-built in 2016. 4. Is longer than 50m. 5. Can be seen from another similar structure. 6. Hosts a yearly dinner. 7. Was originally built for another use case, but has not served that use case for a number of years.

Answer: Shorncliffe Pier, Brisbane, Australia

AA-Omniscience

A private subset of 600 factual questions, balanced across 6 domains. The grader scores accuracy. The score shows how much search adds to the model's internal knowledge.

Representative example

In Karen Russell's short story "St. Lucy's Home for Girls Raised by Wolves," from which locality was the three-piece jazz band hired for the Debutante Ball?

Answer: West Toowoomba

The examples are representative problem types, not samples from the benchmarks.

方法論と詳細情報

Search API の指標は公開ベンチマークの実行結果を集計したもので、利用できる場合はモデルのみのベースラインも含みます。採点、レイテンシ、費用、ベンチマークの網羅範囲の詳細は方法論ページをご覧ください。

よくある質問

Artificial Analysis がベンチマークした 7 社の Search API プロバイダーのうち、Parallel Search(advanced) が 75 で Artificial Analysis Search Index の首位です。

タスク単位では Keenable Search(realtime) が最速で、タスクあたりの平均時間は 16.8s です(モデル時間と検索時間の合計)。検索クエリ単位では Keenable Search(realtime) が最速で、検索クエリあたりの平均時間は 0.34s です。

Parallel Search(turbo) は実測の検索費用が最も低く、ベンチマークタスク 1,000 件あたり $13.64 です。

検索による実測の品質向上が最も大きいのは Parallel Search(advanced) で、検索を使わない同じモデル(モデルのみのベースライン)と比べて Artificial Analysis Search Index が 42 上昇します。

Search API の回答品質は、AA-Omniscience、BrowseComp、DeepSearchQA などの公開ベンチマークを集計して算出します。いずれも、検索を用いて情報を見つけ、抽出し、検証するモデルの能力を測定します。

最適なプロバイダーは優先事項によって変わります。回答の精度を比べるには品質のグラフを、品質と検索・モデル費用のバランスを取るには費用のグラフを、リアルタイム用途にはレイテンシのグラフをご利用ください。トレードオフの散布図では、品質と費用、品質とレイテンシのフロンティア上にあるプロバイダーが分かります。 方法論の全文を見る