Artificial Analysis Search Index: Search API 벤치마크 및 리더보드

Search API는 AI 에이전트가 웹에서 정보를 가져와 최신 출처에 근거한 답변을 만들 수 있게 해줍니다. 딥 리서치, 코딩, 일반 지식 업무에 이르기까지 다양한 에이전트 애플리케이션에서 사실성과 전반적인 결과 품질을 높이는 데 사용됩니다. Search API는 검색 대상이 되는 웹 인덱스, 검색 결과를 모델에 제시하는 방식, 그리고 비용·속도·검색 품질 사이의 절충 등 여러 중요한 지점에서 차이가 있습니다. Artificial Analysis는 7개 제공업체의 Search API 제품 11개를 벤치마크하여, 개발자가 다음 AI 에이전트에 사용할 검색 도구를 고를 때 더 나은 판단을 내릴 수 있도록 돕습니다.

벤치마크 정의, 채점 방식, 데이터 처리에 대한 자세한 내용은 방법론 페이지를 참고하세요.

최신 데이터: Aug 17, 2026.

Stirrup Stirrup agent
Benchmark tasks
Candidate model
GPT-5.6 Luna (medium)
turn 0 / 25
 
Search provider
 
web_search  
native payload, up to 10 results
Grading
Grader
avg F1 0.79
 
Search API providers
Constants
Same tasks, same model and grader (GPT-5.6 Luna, medium), same Stirrup harness (web_search · web_fetch · finish, max 25 turns). Only the search provider varies.

주요 내용

Search API 최고 결과 (대상 벤치마크: 3)
작업당 평균 모델 비용 및 Search API 비용 (USD) · Lower is better
작업당 평균 모델 시간 및 검색 시간 · Lower is better

결과

공개 벤치마크 전반에서 검색을 사용했을 때의 답변 품질입니다.

Artificial Analysis Search Index

DeepSearchQA F1, BrowseComp 정확도, AA-Omniscience 정확도를 동일 가중치로 평균한 값 · Higher is better
모델 단독 (33)

The equal-weighted mean of each Search API provider's score on DeepSearchQA (average F1), BrowseComp (exact-answer accuracy), and AA-Omniscience (accuracy), shown on a 0-100 scale. Every result runs the same candidate answer model, harness, and settings. The only variable is the Search API provider.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

비용

벤치마크 작업당 및 검색 쿼리당 후보 모델 비용과 Search API 비용입니다.

작업당 비용

작업당 평균 후보 모델 비용 및 Search API 비용 (USD) · Lower is better
모델 단독 ($0.0029)

Total cost of one benchmark task, split into the candidate answer model's tokens (input, cached, reasoning, and output) and the Search API provider's list price for the searches the model chose to run. The model-only baseline runs no searches, so its cost is token cost alone. Model cost can differ per provider due to differences in returned payloads for a given query, increased numbers of searches performed (due to not getting the right results) as well as increased reasoning token usage when search results are of lower quality.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

Artificial Analysis Search Index 대비 작업당 비용

Artificial Analysis Search Index와 벤치마크 작업당 총 모델·검색 비용 (USD)
Most attractive quadrant
Pareto line

Total cost of one benchmark task, split into the candidate answer model's tokens (input, cached, reasoning, and output) and the Search API provider's list price for the searches the model chose to run. The model-only baseline runs no searches, so its cost is token cost alone. Model cost can differ per provider due to differences in returned payloads for a given query, increased numbers of searches performed (due to not getting the right results) as well as increased reasoning token usage when search results are of lower quality.

지연 시간

벤치마크 작업당 및 검색 쿼리당 모델 시간과 검색 시간입니다.

작업당 시간

작업당 평균 추정 모델 시간 및 실측 검색 시간 · Lower is better
모델 단독 (15.9s)

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

Artificial Analysis Search Index 대비 작업당 시간

Artificial Analysis Search Index와 벤치마크 작업당 평균 모델·검색 시간 (초)
Most attractive quadrant
Pareto line

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

리더보드 세부 정보

정렬 가능한 Search API 공개 벤치마크 행과 벤치마크별 세부 내역입니다.

Search API 공개 리더보드

제공업체 단위의 Search API 품질, 지연 시간, 비용과 벤치마크별 세부 내역입니다.
12개 행
제공업체
Parallel Search (advanced) logoParallel Search (advanced)
75
42
81
77
67
$47.93
$35.58
37.5s
Exa Search (auto) logoExa Search (auto)
74
41
78
74
70
$65.57
$61.58
27.8s
Firecrawl Search logoFirecrawl Search
73
40
74
74
73
$30.48
$44.94
56.9s
Parallel Search (basic) logoParallel Search (basic)
73
40
79
73
68
$45.14
$69.69
22.2s
Exa Search (fast) logoExa Search (fast)
68
35
76
61
69
$78.11
$80.53
22.6s
You.com Search logoYou.com Search
68
35
63
74
66
$68.93
$58.28
35.8s
Parallel Search (turbo) logoParallel Search (turbo)
67
34
70
75
56
$13.64
$46.57
20.7s
Keenable Search (pro) logoKeenable Search (pro)
67
34
70
66
65
$23.69
$73.20
24.8s
Keenable Search (realtime) logoKeenable Search (realtime)
67
34
70
64
66
$26.87
$63.67
16.8s
Tavily Search (basic) logoTavily Search (basic)
66
33
74
59
64
$126
$66.58
38.6s
Brave Search logoBrave Search
65
32
62
67
65
$71.61
$74.52
28.9s
모델 단독
33
기준선
45
17
38
$2.92
15.9s

Search $/1k and Model $/1k are the average cost of 1,000 benchmark tasks broken down into search and model token costs. The model-only baseline runs no searches.

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

예시 작업

Artificial Analysis Search Index를 구성하는 각 벤치마크의 대표적인 예시 작업입니다.

DeepSearchQA

Broad research questions that need many searches. Answers are lists of items. An LLM grader scores each answer with an F1 score over the answer items. The full eval split has 900 tasks.

Representative example

I've completed the Desert Treasure quest on Old School Runescape, can you list all the names of the spells I can use to teleport into the wilderness that require more than four runes?

Answer: Dareeyak Teleport, Ghorrock Teleport

BrowseComp

Hard-to-find facts that need multi-hop browsing. We use a hard 200-sample subset from the full evaluation pool. The grader checks for an exact answer.

Representative example

I'm looking for the name and location of a structure which fulfills the following criteria: 1. Located in Eastern Australia. 2. Can be visited on foot. 3. Was re-built in 2016. 4. Is longer than 50m. 5. Can be seen from another similar structure. 6. Hosts a yearly dinner. 7. Was originally built for another use case, but has not served that use case for a number of years.

Answer: Shorncliffe Pier, Brisbane, Australia

AA-Omniscience

A private subset of 600 factual questions, balanced across 6 domains. The grader scores accuracy. The score shows how much search adds to the model's internal knowledge.

Representative example

In Karen Russell's short story "St. Lucy's Home for Girls Raised by Wolves," from which locality was the three-piece jazz band hired for the Debutante Ball?

Answer: West Toowoomba

The examples are representative problem types, not samples from the benchmarks.

방법론 및 추가 정보

Search API 지표는 공개 벤치마크 실행 결과를 집계한 것이며, 가능한 경우 모델 단독 기준선도 포함합니다. 채점, 지연 시간, 비용, 벤치마크 적용 범위에 대한 자세한 내용은 방법론 페이지를 참고하세요.

자주 묻는 질문

Artificial Analysis가 벤치마크한 Search API 제공업체 7곳 가운데 Parallel Search (advanced)가 75로 Artificial Analysis Search Index 1위입니다.

작업 기준으로는 Keenable Search (realtime)가 가장 빠르며, 작업당 평균 시간은 16.8s입니다(모델 시간과 검색 시간의 합). 개별 검색 쿼리 기준으로는 Keenable Search (realtime)가 가장 빠르며, 검색 쿼리당 평균 시간은 0.34s입니다.

Parallel Search (turbo)의 실측 검색 비용이 가장 낮으며, 벤치마크 작업 1,000건당 $13.64입니다.

검색으로 측정된 품질 향상이 가장 큰 곳은 Parallel Search (advanced)로, 검색을 쓰지 않은 동일 모델(모델 단독 기준선) 대비 Artificial Analysis Search Index가 42 상승합니다.

Search API의 답변 품질은 AA-Omniscience, BrowseComp 및 DeepSearchQA 등 공개 벤치마크를 집계해 산출합니다. 각 벤치마크는 검색을 활용해 정보를 찾고, 추출하고, 검증하는 모델의 능력을 측정합니다.

가장 적합한 제공업체는 우선순위에 따라 달라집니다. 답변 정확도를 비교하려면 품질 차트를, 품질과 검색·모델 비용의 균형을 맞추려면 비용 차트를, 실시간 사용 사례에는 지연 시간 차트를 활용하세요. 절충 관계를 보여주는 산점도에서는 품질-비용 및 품질-지연 시간 프런티어에 있는 제공업체를 확인할 수 있습니다. 전체 방법론 보기