Artificial Analysis Search Index: Search API 벤치마크 및 리더보드

Search API는 AI 에이전트가 웹에서 정보를 가져와 최신 출처에 근거한 답변을 만들 수 있게 해줍니다. 딥 리서치, 코딩, 일반 지식 업무에 이르기까지 다양한 에이전트 애플리케이션에서 사실성과 전반적인 결과 품질을 높이는 데 사용됩니다. Search API는 검색 대상이 되는 웹 인덱스, 검색 결과를 모델에 제시하는 방식, 그리고 비용·속도·검색 품질 사이의 절충 등 여러 중요한 지점에서 차이가 있습니다. Artificial Analysis는 12개 제공업체의 Search API 제품 25개를 벤치마크하여, 개발자가 다음 AI 에이전트에 사용할 검색 도구를 고를 때 더 나은 판단을 내릴 수 있도록 돕습니다.

벤치마크 정의, 채점 방식, 데이터 처리에 대한 자세한 내용은 방법론 페이지를 참고하세요.

최신 데이터: Sep 22, 2026.

주요 내용

Search API 최고 결과 (대상 벤치마크: 3)
작업당 평균 모델 비용 및 Search API 비용 (USD) · Lower is better
작업당 평균 모델 시간 및 검색 시간 · Lower is better

결과

공개 벤치마크 전반에서 검색을 사용했을 때의 답변 품질입니다.

Artificial Analysis Search Index

DeepSearchQA F1, BrowseComp 정확도, AA-Omniscience 정확도를 동일 가중치로 평균한 값 · Higher is better
모델 단독 (33)

비용

벤치마크 작업당 및 검색 쿼리당 후보 모델 비용과 Search API 비용입니다.

작업당 비용

작업당 평균 후보 모델 비용 및 Search API 비용 (USD) · Lower is better
모델 단독 ($0.0029)

Artificial Analysis Search Index 대비 작업당 비용

Artificial Analysis Search Index와 벤치마크 작업당 총 모델·검색 비용 (USD)
Most attractive quadrant
Pareto line

지연 시간

벤치마크 작업당 및 검색 쿼리당 모델 시간과 검색 시간입니다.

작업당 시간

작업당 평균 추정 모델 시간 및 실측 검색 시간 · Lower is better
모델 단독 (22.5s)

Artificial Analysis Search Index 대비 작업당 시간

Artificial Analysis Search Index와 벤치마크 작업당 평균 모델·검색 시간 (초)
Most attractive quadrant
Pareto line

토큰

벤치마크 작업당 후보 모델의 입력 및 출력 토큰 수와 전체 토큰 사용량 대비 품질입니다.

작업당 총 토큰 수

벤치마크 작업당 후보 모델의 평균 입력, 추론 및 답변 토큰 수 · Lower is better
모델 단독 (2.6k)

Artificial Analysis Search Index 대비 작업당 총 토큰 수

Artificial Analysis Search Index와 벤치마크 작업당 후보 모델의 총 토큰 수
Most attractive quadrant
Pareto line

검색

벤치마크 작업당 후보 모델이 실행하는 검색 쿼리 수입니다.

작업당 검색 쿼리 수

벤치마크 작업당 후보 모델이 실행하는 평균 검색 쿼리 수

리더보드 세부 정보

정렬 가능한 Search API 공개 벤치마크 행과 벤치마크별 세부 내역입니다.

Search API 공개 리더보드

제공업체 단위의 Search API 품질, 지연 시간, 비용과 벤치마크별 세부 내역입니다.
13개 행
제공업체
Perplexity Search (medium) logoPerplexity Search (medium)
80
47
81
87
72
$62.30
$29.09
28.5s
Perplexity Search (high) logoPerplexity Search (high)
79
46
81
86
71
$56.93
$34.44
29.3s
Octen Search (highlights) logoOcten Search (highlights)
77
44
80
86
66
$9.07
$49.15
16.9s
Parallel Search (advanced) logoParallel Search (advanced)
75
42
81
77
67
$47.93
$35.58
42.5s
Brave Search (LLM context) logoBrave Search (LLM context)
75
42
78
77
69
$61.96
$67.57
23.9s
You.com Search (highlights) logoYou.com Search (highlights)
74
41
77
77
69
$47.54
$69.94
20.4s
Nimble Search (standard) logoNimble Search (standard)
74
41
76
75
71
$46.14
$59.91
44.4s
Exa Search (auto) logoExa Search (auto)
74
41
78
74
70
$65.57
$61.58
32.9s
Firecrawl Search logoFirecrawl Search
73
40
74
74
73
$30.48
$44.94
62.7s
Parallel Search (basic) logoParallel Search (basic)
73
40
79
73
68
$45.14
$69.69
27.5s
TinyFish Search (web) logoTinyFish Search (web)
71
38
67
75
71
$0
$34.55
59.8s
Exa Search (instant) logoExa Search (instant)
68
35
73
65
67
$72.23
$78.69
19.9s
모델 단독
33
기준선
45
17
38
—
$2.92
22.5s

예시 작업

Artificial Analysis Search Index를 구성하는 각 벤치마크의 대표적인 예시 작업입니다.

DeepSearchQA

Broad research questions that need many searches. Answers are lists of items. An LLM grader scores each answer with an F1 score over the answer items. The full eval split has 900 tasks.

Representative example

I've completed the Desert Treasure quest on Old School Runescape, can you list all the names of the spells I can use to teleport into the wilderness that require more than four runes?

Answer: Dareeyak Teleport, Ghorrock Teleport

BrowseComp

Hard-to-find facts that need multi-hop browsing. We use a hard 200-sample subset from the full evaluation pool. The grader checks for an exact answer.

Representative example

I'm looking for the name and location of a structure which fulfills the following criteria: 1. Located in Eastern Australia. 2. Can be visited on foot. 3. Was re-built in 2016. 4. Is longer than 50m. 5. Can be seen from another similar structure. 6. Hosts a yearly dinner. 7. Was originally built for another use case, but has not served that use case for a number of years.

Answer: Shorncliffe Pier, Brisbane, Australia

AA-Omniscience

A private subset of 600 factual questions, balanced across 6 domains. The grader scores accuracy. The score shows how much search adds to the model's internal knowledge.

Representative example

In Karen Russell's short story "St. Lucy's Home for Girls Raised by Wolves," from which locality was the three-piece jazz band hired for the Debutante Ball?

Answer: West Toowoomba

The examples are representative problem types, not samples from the benchmarks.

방법론 및 추가 정보

Search API 지표는 공개 벤치마크 실행 결과를 집계한 것이며, 가능한 경우 모델 단독 기준선도 포함합니다. 채점, 지연 시간, 비용, 벤치마크 적용 범위에 대한 자세한 내용은 방법론 페이지를 참고하세요.

자주 묻는 질문