Artificial Analysis Search Index: Search-API-Benchmark und Bestenliste

Search APIs ermöglichen es KI-Agenten, Informationen aus dem Web abzurufen und ihre Antworten auf aktuelle Quellen zu stützen. Sie kommen in einer Vielzahl agentischer Anwendungen zum Einsatz – von Deep Research über die Programmierung bis hin zur allgemeinen Wissensarbeit –, um Faktentreue und Ergebnisqualität zu verbessern. Search APIs unterscheiden sich in mehreren wichtigen Punkten: im zugrunde liegenden Web-Index, den sie durchsuchen, in der Art, wie Suchergebnisse dem Modell präsentiert werden, und in den Kompromissen, die sie zwischen Kosten, Geschwindigkeit und Retrieval-Qualität eingehen. Wir benchmarken 11 Search-API-Produkte von 7 Anbietern, damit Entwickler fundiertere Entscheidungen treffen können, wenn sie ein Suchwerkzeug für ihren nächsten KI-Agenten auswählen.

Benchmark-Definitionen, Bewertung und Details zur Datenverarbeitung finden Sie auf der Methodikseite.

Neueste Daten: Aug 17, 2026.

Stirrup Stirrup agent
Benchmark tasks
Candidate model
GPT-5.6 Luna (medium)
turn 0 / 25
 
Search provider
 
web_search  
native payload, up to 10 results
Grading
Grader
avg F1 0.79
 
Search API providers
Constants
Same tasks, same model and grader (GPT-5.6 Luna, medium), same Stirrup harness (web_search · web_fetch · finish, max 25 turns). Only the search provider varies.

Wichtigste Ergebnisse

Bestes Search-API-Ergebnis über 3 Benchmarks
Durchschnittliche Modell- und Search-API-Kosten pro Aufgabe (USD) · Lower is better
Durchschnittliche Modell- und Suchzeit pro Aufgabe · Lower is better

Ergebnisse

Antwortqualität mit Suche über öffentliche Benchmarks hinweg.

Artificial Analysis Search Index

Gleichgewichteter Mittelwert aus DeepSearchQA F1, BrowseComp-Genauigkeit und AA-Omniscience-Genauigkeit · Higher is better
Nur Modell (33)

The equal-weighted mean of each Search API provider's score on DeepSearchQA (average F1), BrowseComp (exact-answer accuracy), and AA-Omniscience (accuracy), shown on a 0-100 scale. Every result runs the same candidate answer model, harness, and settings. The only variable is the Search API provider.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

Kosten

Kosten des Kandidatenmodells und der Search API pro Benchmark-Aufgabe und pro Suchanfrage.

Kosten pro Aufgabe

Durchschnittliche Kosten des Kandidatenmodells und der Search API pro Aufgabe (USD) · Lower is better
Nur Modell ($0.0029)

Total cost of one benchmark task, split into the candidate answer model's tokens (input, cached, reasoning, and output) and the Search API provider's list price for the searches the model chose to run. The model-only baseline runs no searches, so its cost is token cost alone. Model cost can differ per provider due to differences in returned payloads for a given query, increased numbers of searches performed (due to not getting the right results) as well as increased reasoning token usage when search results are of lower quality.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

Artificial Analysis Search Index vs. Kosten pro Aufgabe

Artificial Analysis Search Index vs. Gesamtkosten für Modell und Suche pro Benchmark-Aufgabe (USD)
Most attractive quadrant
Pareto line

Total cost of one benchmark task, split into the candidate answer model's tokens (input, cached, reasoning, and output) and the Search API provider's list price for the searches the model chose to run. The model-only baseline runs no searches, so its cost is token cost alone. Model cost can differ per provider due to differences in returned payloads for a given query, increased numbers of searches performed (due to not getting the right results) as well as increased reasoning token usage when search results are of lower quality.

Latenz

Modellzeit und Suchzeit pro Benchmark-Aufgabe und pro Suchanfrage.

Zeit pro Aufgabe

Durchschnittliche abgeleitete Modellzeit und gemessene Suchzeit pro Aufgabe · Lower is better
Nur Modell (15.9s)

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

The dashed line represents the candidate answer model, GPT-5.6 Luna (medium), run with no search tools and is considered the baseline. See the Search API methodology for the full harness settings.

Artificial Analysis Search Index vs. Zeit pro Aufgabe

Artificial Analysis Search Index vs. durchschnittliche Modell- und Suchzeit pro Benchmark-Aufgabe (Sekunden)
Most attractive quadrant
Pareto line

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

Details zur Bestenliste

Sortierbare öffentliche Search-API-Benchmark-Zeilen und Aufschlüsselung je Benchmark.

Öffentliche Search-API-Bestenliste

Search-API-Qualität, -Latenz und -Kosten nach Anbieter, mit Aufschlüsselung je Benchmark.
12 Zeilen
Anbieter
Parallel Search (advanced) logoParallel Search (advanced)
75
42
81
77
67
$47.93
$35.58
37.5s
Exa Search (auto) logoExa Search (auto)
74
41
78
74
70
$65.57
$61.58
27.8s
Firecrawl Search logoFirecrawl Search
73
40
74
74
73
$30.48
$44.94
56.9s
Parallel Search (basic) logoParallel Search (basic)
73
40
79
73
68
$45.14
$69.69
22.2s
Exa Search (fast) logoExa Search (fast)
68
35
76
61
69
$78.11
$80.53
22.6s
You.com Search logoYou.com Search
68
35
63
74
66
$68.93
$58.28
35.8s
Parallel Search (turbo) logoParallel Search (turbo)
67
34
70
75
56
$13.64
$46.57
20.7s
Keenable Search (pro) logoKeenable Search (pro)
67
34
70
66
65
$23.69
$73.20
24.8s
Keenable Search (realtime) logoKeenable Search (realtime)
67
34
70
64
66
$26.87
$63.67
16.8s
Tavily Search (basic) logoTavily Search (basic)
66
33
74
59
64
$126
$66.58
38.6s
Brave Search logoBrave Search
65
32
62
67
65
$71.61
$74.52
28.9s
Nur Modell
33
Baseline
45
17
38
$2.92
15.9s

Search $/1k and Model $/1k are the average cost of 1,000 benchmark tasks broken down into search and model token costs. The model-only baseline runs no searches.

The sum of model time and search time for one benchmark task. Model time is derived: answer and reasoning tokens per task divided by the model's canonical answer output speed. Search time is the measured time spent in web_search calls. A provider can be fast per call and still add more total time if the model searches against it more often.

Beispielaufgaben

Repräsentative Beispielaufgaben für jeden Benchmark hinter dem Artificial Analysis Search Index.

DeepSearchQA

Broad research questions that need many searches. Answers are lists of items. An LLM grader scores each answer with an F1 score over the answer items. The full eval split has 900 tasks.

Representative example

I've completed the Desert Treasure quest on Old School Runescape, can you list all the names of the spells I can use to teleport into the wilderness that require more than four runes?

Answer: Dareeyak Teleport, Ghorrock Teleport

BrowseComp

Hard-to-find facts that need multi-hop browsing. We use a hard 200-sample subset from the full evaluation pool. The grader checks for an exact answer.

Representative example

I'm looking for the name and location of a structure which fulfills the following criteria: 1. Located in Eastern Australia. 2. Can be visited on foot. 3. Was re-built in 2016. 4. Is longer than 50m. 5. Can be seen from another similar structure. 6. Hosts a yearly dinner. 7. Was originally built for another use case, but has not served that use case for a number of years.

Answer: Shorncliffe Pier, Brisbane, Australia

AA-Omniscience

A private subset of 600 factual questions, balanced across 6 domains. The grader scores accuracy. The score shows how much search adds to the model's internal knowledge.

Representative example

In Karen Russell's short story "St. Lucy's Home for Girls Raised by Wolves," from which locality was the three-piece jazz band hired for the Debutante Ball?

Answer: West Toowoomba

The examples are representative problem types, not samples from the benchmarks.

Methodik und weitere Informationen

Search-API-Kennzahlen werden aus öffentlichen Benchmark-Durchläufen aggregiert und enthalten, soweit verfügbar, Baselines mit reinem Modellbetrieb. Details zu Bewertung, Latenz, Kosten und Benchmark-Abdeckung finden Sie auf der Methodikseite.

Häufig gestellte Fragen

Parallel Search (advanced) führt den Artificial Analysis Search Index mit 75 an, gemessen über 7 von Artificial Analysis untersuchte Search-API-Anbieter.

Pro Aufgabe ist Keenable Search (realtime) am schnellsten, mit einer durchschnittlichen Zeit pro Aufgabe von 16.8s (Modellzeit plus Suchzeit). Pro einzelner Suchanfrage ist Keenable Search (realtime) am schnellsten, mit einer durchschnittlichen Zeit pro Suchanfrage von 0.34s.

Parallel Search (turbo) hat mit $13.64 pro 1.000 Benchmark-Aufgaben die niedrigsten gemessenen Suchkosten.

Den größten gemessenen Qualitätsgewinn bringt die Suche bei Parallel Search (advanced): Der Artificial Analysis Search Index steigt um 42 gegenüber demselben Modell ohne Suche (seiner Baseline mit reinem Modellbetrieb).

Die Antwortqualität von Search APIs wird aus öffentlichen Benchmarks aggregiert, darunter AA-Omniscience, BrowseComp und DeepSearchQA. Jeder misst die Fähigkeit eines Modells, Informationen mithilfe von Suche zu finden, zu extrahieren oder zu verifizieren.

Der beste Anbieter hängt von Ihren Prioritäten ab. Nutzen Sie die Qualitätsdiagramme, um die Antwortgenauigkeit zu vergleichen, die Kostendiagramme, um Qualität gegen Such- und Modellkosten abzuwägen, und die Latenzdiagramme für Echtzeit-Anwendungsfälle. Die Streudiagramme zu den Zielkonflikten heben die Anbieter auf der Qualität-Kosten- und der Qualität-Latenz-Front hervor. Vollständige Methodik ansehen