Speech to Speech Benchmarking Methodology

Overview

Our current Speech to Speech benchmarking evaluates native audio models - models that support native audio input and output - across four complementary dimensions: speech reasoning, conversational dynamics, agentic performance and human preference in live conversations.

Artificial Analysis Speech to Speech Index

The Artificial Analysis Speech to Speech Index is an equal-weighted score for native audio models. It combines results across Speech Reasoning, Agentic Performance, Arena Preference, and Task Success Rate. Models must have valid results for all four components to appear in the index.

Current weighting is equal across the four components: 25% Speech Reasoning (Big Bench Audio), 25% Agentic Performance (๐œ-Voice), 25% Arena Preference (Arena Score), and 25% Task Success Rate.

Speech Reasoning

Overview

Our speech reasoning benchmark evaluates the ability of native audio models to answer reasoning-based questions.

Native audio models are provided with an input audio file and are expected to generate an output audio. The judge model is provided with the candidate answer, official answer and original question as context and is prompted to label the candidate answer as correct or incorrect.

Dataset: Big Bench Audio

The emergence of native audio-to-audio models offers exciting opportunities to increase voice agent capabilities and simplify workflows. However, it's crucial to evaluate whether this simplification comes at the cost of model performance or introduces other trade-offs.

To help answer this question we've released Big Bench Audio, a new dataset for benchmarking the performance of native audio models.

Big Bench Audio contains 1,000 audio files representing questions designed to test the intelligence of models. The questions are based on four categories of the Big Bench Hard dataset (250 questions each) and were generated using 23 synthetic voices from top-ranked text to speech models in the Artificial Analysis Text to Speech Arena.

Question Categories

Formal Fallacies (250 questions): Determine whether an argument presented informally can be logically deduced from the provided context.

Example: First of all, everyone who is a close friend of Glenna is a close friend of Tamara, too. Next, whoever is neither a half-sister of Deborah nor a workmate of Nila is a close friend of Glenna. Hence, whoever is none of this: a half-sister of Deborah or workmate of Nila, is a close friend of Tamara. Is the argument deductively valid or invalid?

Navigate (250 questions): Determine whether a series of navigation steps returns the agent to the starting point.

Example: If you follow these instructions, do you return to the starting point? Take 10 steps. Turn around. Take 4 steps. Take 6 steps. Turn around.

Object Counting (250 questions): Count the number of a specific item class given a collection of possessions.

Example: I have three blackberries, two strawberries, an apple, three oranges, a nectarine, a grape, a peach, a banana, and a plum. How many fruits do I have?

Web of Lies (250 questions): Evaluate the truth value of a Boolean function expressed as a natural-language word problem.

Example: Leda tells the truth. Alexis says Leda lies. Sal says Alexis lies. Phoebe says Sal tells the truth. Gwenn says Phoebe tells the truth. Does Gwenn tell the truth?

To enable evaluation of the tradeoffs associated with using native speech to speech models we test multiple different configurations on Big Bench Audio. To learn more about Big Bench Audio, review the article or download the dataset yourself.

Conversational Dynamics

Overview

Our conversational dynamics benchmark evaluates the ability of native audio models to handle realistic conversational behaviors - the kinds of interactions that occur naturally in human conversation but are challenging for speech models to manage correctly.

This benchmark is implemented by Artificial Analysis based on a subset of Full Duplex Bench v1 (Lin et al., 2025) and Full Duplex Bench v1.5 (Lin et al., 2025), benchmarks that systematically evaluate key interactive behaviors of full duplex spoken dialogue models.

Metrics

From Full Duplex Bench v1:

  • Pause Handling: Percentage of samples where the model correctly does not interrupt during a user's natural pause. Evaluates whether the model recognizes that the speaker still holds the floor.
  • Turn Taking: Percentage of samples where the model correctly takes the conversational turn when appropriate. Measures the model's ability to detect turn boundaries and respond promptly.

From Full Duplex Bench v1.5:

  • User Interruption Handling: Percentage of samples where the model correctly addresses the user's interruption - responding to questions or changes in topic raised mid-conversation.
  • Backchannel Handling: Percentage of samples where the model correctly continues its response when a backchannel such as "yeah", "alright", or "mm-hmm" is played, rather than treating it as a new turn.

Agentic Performance (๐œ-Voice)

Overview

Our Agentic Performance benchmark evaluates the ability of Speech to Speech models to complete realistic customer service tasks end-to-end. This benchmark measures multi-turn instruction following, the ability to support a simulated customer through a complete interaction, and successful tool use against simulated customer service systems.

This benchmark is implemented by Artificial Analysis based on ๐œ-Voice (Ray, Dhandhania, Barres & Narasimhan, 2026), a benchmark by Sierra that evaluates full duplex voice agents on grounded customer service tasks across real-world domains.

Metric

  • Task Completion (pass@1): Proportion of scenarios where the model correctly resolves the customer's issue. Each score is the mean of three independent trials. The tested model is prompted as a customer support agent with access to domain-specific tools and a policy document. Each scenario has a single valid database end-state; evaluation compares the final state against this ground truth.

We evaluate across three domains using the base task set:

  • Airline (50 scenarios): e.g., changing a flight, rebooking under policy constraints
  • Retail (114 scenarios): e.g., disputing a charge, processing a return
  • Telecom (114 scenarios): e.g., resolving a billing issue, troubleshooting a service problem

Voice Personas

Customer voices are generated using ElevenLabs, based on prompts adapted from the publicly available Sierra ๐œ-Voice implementation. We use two control personas representing standard English speakers, and five regular personas representing diverse accents and speaker profiles.

Speech Agent Arena

Overview

The Speech Agent Arena is a blind preference benchmark that evaluates which native audio model participants prefer in live voice conversations. It complements our automated benchmarks by measuring the end-to-end conversational experience on realistic tasks with human participants.

In each round, a participant receives one scenario, completes it separately with two hidden models, and records a forced overall preference after both calls. Participants also answer diagnostic questions, which we record separately from the overall preference vote.

Metrics

  • Preference Elo: Preference Elo is calculated separately across all scenarios, agentic scenarios, and non-agentic scenarios using Bradleyโ€“Terry maximum likelihood. Approximate 95% confidence intervals use the same Hessian-based method as the TTS Arena, with GPT Realtime 1.5 anchoring the scale at 1000 Elo.
  • Task Success Rate: The Task Success Rate is the percentage of eligible conversations where the model made the correct final task-completing tool call or calls. Eligible conversations equal successes plus model failures; participant deviations and unverifiable cases are excluded before calculation.

Dataset / Evaluation Setup

The Arena includes 35 scenarios: 15 agentic scenarios with tool calls and 20 non-agentic scenarios without tools.

  • Agentic (15 scenarios): Models receive tools relevant to completing the assigned task.

    Examples:

    • Book a new-patient dental check-up for Tuesday at 9:30am, or the earliest available morning if that time is taken.
    • Order two different pizzas and one side for delivery, keeping the total including delivery under $45.
  • Non-agentic (20 scenarios): Models complete the conversation without tools, using information supplied in the system prompt.

    Examples:

    • Ask about Sunday pool hours, family entry prices and towel availability.
    • Ask about beginner yoga classes, class and membership prices, and what to bring.

See the Speech Agent Arena Overview for a worked Arena round, model inputs, tool schemas and example conversations.

Task Success Rate

Step 1: Participant eligibility. Judge 1 receives the assigned scenario, transcript and required final-call definitions. It checks that the participant attempted the assigned task, stayed materially within it, communicated a final request or accepted outcome that can be assessed, and gave the model a reasonable opportunity to act. Only eligible conversations continue to Step 2.

Step 2a: Required final calls. For eligible conversations, Judge 2 receives the required final calls and complete chronological tool trace. It checks that every required final call was made with the correct tool and call count, with no unintended final actions. Supporting calls and end_call are not scored as task completion.

Step 2b: Final-call arguments. The same Judge 2 assessment checks that every schema-required argument was supplied and matched the participantโ€™s final spoken request and conversation context. Operationally equivalent wording and harmless transcription variations are allowed. A recovered failed attempt can pass; missing calls, wrong arguments or extra final actions fail.

Publication rules

  • Overall results are published for a model with at least 100 appearances and a 95% confidence interval half-width no greater than 75.
  • Participant recordings and transcripts are not published. Examples will appear only once recording consent is confirmed and any personal information has been reviewed and redacted.

We report a 95% confidence interval for each model's Task Success Rate using the Wilson score interval. Participant deviations and unverifiable conversations are excluded before calculation.

Speech to Speech Index Integration

For use only in the Artificial Analysis Speech to Speech Index:

Arena Score: Arena Score converts a model's Elo into expected preference against an 800 Elo baseline, using an Elo frozen at the point the model became eligible for publication.

Arena Score = 100 / (1 + 10^((800 โˆ’ frozen Elo) / 400))

Price

  • Price per Hour of Input Audio: Total cost (USD) of audio included in the request / message sent to the API.
  • Price per Hour of Output Audio: Total cost (USD) of audio generated by the model (received from the API).

Speed

  • Time to First Audio (TTFA): Average number of seconds required to generate the first token of audio output, measured across the Big Bench Audio question set. TTFA is a critical indicator of perceived responsiveness in voice agent applications.

Version History

Artificial Analysis Speech to Speech Index v2.0

August 2026 to Present

  • Replaced Conversational Dynamics (Full Duplex Bench) with Task Success Rate in the index.
  • Equal weighting across four components: 25% Speech Reasoning (Big Bench Audio), 25% Agentic Performance (๐œ-Voice), 25% Arena Preference (Arena Score), and 25% Task Success Rate.

Artificial Analysis Speech to Speech Index v1.1

August 2026

  • Added Speech Agent Arena to the Artificial Analysis Speech to Speech Index.
  • Equal weighting across four datasets: 25% Speech Reasoning (Big Bench Audio), 25% Conversational Dynamics (Full Duplex Bench), 25% Agentic Performance (๐œ-Voice), and 25% Speech Agent Arena.

Artificial Analysis Speech to Speech Index v1.0

June 2026 to August 2026

  • Launched the Artificial Analysis Speech to Speech Index, a weighted average score requiring Speech Reasoning, Conversational Dynamics and Agentic Performance results.
  • Equal weighting across the three datasets: 33.3% Speech Reasoning (Big Bench Audio), 33.3% Conversational Dynamics (Full Duplex Bench), and 33.3% Agentic Performance (๐œ-Voice).

Big Bench Audio (BBA) v1.2

May 2026 to Present

  • Updated the judge model and grader harness to more reliably recognize correct final answers in verbose responses, including cases where the model discusses incorrect alternatives before giving the correct answer.

Big Bench Audio (BBA) v1.1

March 2026 to May 2026

  • Accuracy measured as the number of correct answers out of 1,000, including questions where the model did not answer.
  • Claude Sonnet 4.6 used as the judge model.

Big Bench Audio (BBA) v1.0

December 2024 to March 2026

  • Accuracy measured as the share of non-error responses answered correctly, so models were not penalized for questions where they did not answer.
  • Claude Sonnet 3.5 used as the judge model.