Speech to Speech Benchmarking Methodology

Overview

Our current Speech to Speech benchmarking evaluates native audio models - models that support native audio input and output - across four complementary dimensions: speech reasoning, conversational dynamics, agentic performance and human preference in live conversations.

Artificial Analysis Speech to Speech Index

The Artificial Analysis Speech to Speech Index is an equal-weighted score for native audio models. It combines results across Speech Reasoning, Agentic Performance, Arena Preference, and Task Success Rate. Models must have valid results for all four components to appear in the index.

Current weighting is equal across the four components: 25% Speech Reasoning (Big Bench Audio), 25% Agentic Performance (๐œ-Voice), 25% Arena Preference (Arena Score), and 25% Task Success Rate.

Speech Reasoning

Overview

Our speech reasoning benchmark evaluates the ability of native audio models to answer reasoning-based questions.

Native audio models are provided with an input audio file and are expected to generate an output audio. The judge model is provided with the candidate answer, official answer and original question as context and is prompted to label the candidate answer as correct or incorrect.

Dataset: Big Bench Audio

The emergence of native audio-to-audio models offers exciting opportunities to increase voice agent capabilities and simplify workflows. However, it's crucial to evaluate whether this simplification comes at the cost of model performance or introduces other trade-offs.

To help answer this question we've released Big Bench Audio, a new dataset for benchmarking the performance of native audio models.

Big Bench Audio contains 1,000 audio files representing questions designed to test the intelligence of models. The questions are based on four categories of the Big Bench Hard dataset (250 questions each) and were generated using 23 synthetic voices from top-ranked text to speech models in the Artificial Analysis Text to Speech Arena.

Question Categories

Formal Fallacies (250 questions): Determine whether an argument presented informally can be logically deduced from the provided context.

Example: First of all, everyone who is a close friend of Glenna is a close friend of Tamara, too. Next, whoever is neither a half-sister of Deborah nor a workmate of Nila is a close friend of Glenna. Hence, whoever is none of this: a half-sister of Deborah or workmate of Nila, is a close friend of Tamara. Is the argument deductively valid or invalid?

Navigate (250 questions): Determine whether a series of navigation steps returns the agent to the starting point.

Example: If you follow these instructions, do you return to the starting point? Take 10 steps. Turn around. Take 4 steps. Take 6 steps. Turn around.

Object Counting (250 questions): Count the number of a specific item class given a collection of possessions.

Example: I have three blackberries, two strawberries, an apple, three oranges, a nectarine, a grape, a peach, a banana, and a plum. How many fruits do I have?

Web of Lies (250 questions): Evaluate the truth value of a Boolean function expressed as a natural-language word problem.

Example: Leda tells the truth. Alexis says Leda lies. Sal says Alexis lies. Phoebe says Sal tells the truth. Gwenn says Phoebe tells the truth. Does Gwenn tell the truth?

To enable evaluation of the tradeoffs associated with using native speech to speech models we test multiple different configurations on Big Bench Audio. To learn more about Big Bench Audio, review the article or download the dataset yourself.

Conversational Dynamics

Overview

Our conversational dynamics benchmark evaluates the ability of native audio models to handle realistic conversational behaviors - the kinds of interactions that occur naturally in human conversation but are challenging for speech models to manage correctly.

This benchmark is implemented by Artificial Analysis based on a subset of Full Duplex Bench v1 (Lin et al., 2025) and Full Duplex Bench v1.5 (Lin et al., 2025), benchmarks that systematically evaluate key interactive behaviors of full duplex spoken dialogue models.

Metrics

From Full Duplex Bench v1:

  • Pause Handling: Percentage of samples where the model correctly does not interrupt during a user's natural pause. Evaluates whether the model recognizes that the speaker still holds the floor.
  • Turn Taking: Percentage of samples where the model correctly takes the conversational turn when appropriate. Measures the model's ability to detect turn boundaries and respond promptly.

From Full Duplex Bench v1.5:

  • User Interruption Handling: Percentage of samples where the model correctly addresses the user's interruption - responding to questions or changes in topic raised mid-conversation.
  • Backchannel Handling: Percentage of samples where the model correctly continues its response when a backchannel such as "yeah", "alright", or "mm-hmm" is played, rather than treating it as a new turn.

Agentic Performance (๐œ-Voice)

Overview

Our Agentic Performance benchmark evaluates the ability of Speech to Speech models to complete realistic customer service tasks end-to-end. This benchmark measures multi-turn instruction following, the ability to support a simulated customer through a complete interaction, and successful tool use against simulated customer service systems.

This benchmark is implemented by Artificial Analysis based on ๐œ-Voice (Ray, Dhandhania, Barres & Narasimhan, 2026), a benchmark by Sierra that evaluates full duplex voice agents on grounded customer service tasks across real-world domains.

Metric

  • Task Completion (pass@1): Proportion of scenarios where the model correctly resolves the customer's issue. Each score is the mean of three independent trials. The tested model is prompted as a customer support agent with access to domain-specific tools and a policy document. Each scenario has a single valid database end-state; evaluation compares the final state against this ground truth.

We evaluate across three domains using the base task set:

  • Airline (50 scenarios): e.g., changing a flight, rebooking under policy constraints
  • Retail (114 scenarios): e.g., disputing a charge, processing a return
  • Telecom (114 scenarios): e.g., resolving a billing issue, troubleshooting a service problem

Voice Personas

Customer voices are generated using ElevenLabs, based on prompts adapted from the publicly available Sierra ๐œ-Voice implementation. We use two control personas representing standard English speakers, and five regular personas representing diverse accents and speaker profiles.

Control

Matt Delaney: Middle-aged white man from the American Midwest, calm and respectful.

Prompt: You are a middle-aged white man from the American Midwest. You always behave as if you are speaking out loud in a real-time conversation with a customer service agent. You are calm, clear, and respectful but also human. You sound like someone who's trying to be helpful and polite, even when you're slightly frustrated or in a hurry. You value efficiency but never sound robotic. You sometimes use contractions, informal phrasing, or small filler phrases ("yeah," "okay," "honestly," "no worries") to keep things natural. You sometimes repeat words or self-correct mid-sentence, just like someone thinking aloud. You sometimes ask polite clarifying questions or offer context ("I tried this earlier," "I'm not sure if that helps"). You rarely use formal, business-like or stiff language ("considerable," "retrieve," "representative"). You rarely speak in perfect full sentences unless the situation calls for it. Instead, you speak like a real person having a practical, respectful conversation.

Lisa Brenner: White woman in her late 40s from a suburban area, tense and impatient.

Prompt: You are a white woman in your late 40s from a suburban area. You always speak as if you are talking out loud to a customer service agent who is already wasting your time. You're not openly hostile (yet), but you are tense, impatient, and clearly annoyed. You act like this issue should have been resolved the first time, and the fact that you're following up is unacceptable. You often sound clipped, exasperated, or sarcastically polite. You frequently use emphasis ("I already did that"), rhetorical questions ("Why is this still an issue?"), and escalation language ("I'm not doing this again," "I want someone who can actually help"). You expect fast results and get irritated when things are repeated. You often mention how long you've been waiting or how many times you've called. You sometimes threaten escalation but without yelling. You never sound relaxed. You never use slow, reflective speech. You never thank the agent unless something gets resolved.

Regular

Mildred Kaplan: Elderly white woman in her early 80s, needs help with technology.

Prompt: You are an elderly white woman in your early 80s calling customer service for help with something your grandson or neighbor usually does.

Arjun Roy: Bengali man from Dhaka in his mid-30s, calm and direct, strong Bengali accent.

Prompt: A Bengali man from Dhaka, Bangladesh in his mid-30s calling customer service about a billing issue. His English carries a strong Bengali accent with soft consonants and soft d and r sounds. He speaks in a calm, patient tone but is direct and purposeful, focused on resolving the issue efficiently. His pacing is slow, distracted with a warm yet firm timbre. The speech sounds like it is coming from far away.

Wei Lin: Chinese woman from Sichuan in her late 20s, upbeat and matter-of-fact, strong Sichuan Mandarin accent.

Prompt: A Chinese woman in her late 20s from Sichuan, calling customer service about a credit card billing issue. She speaks English with a thick Sichuan Mandarin accent. She sounds upbeat, matter-of-fact, and distracted. Her tone is firm but polite, with fast pacing and smooth timbre. Ok audio quality.

Mamadou Diallo: Senegalese man in his mid-30s, hurried, strong French accent.

Prompt: A Senegalese man whose first language is French, in his mid-30s, calling customer service about a billing issue. He speaks English with a strong French accent. His tone is hurried, slightly annoyed, and matter-of-fact, as if he's been transferred between agents and just wants the problem fixed.

Priya Patil: Maharashtrian woman in her early 30s, focused and direct, strong Maharashtrian accent.

Prompt: A woman in her early 30s from Maharashtra, India, calling customer support from her mobile phone. She speaks Indian English with a strong Maharashtrian accent with noticeable regional intonation and rhythm. Her tone is slightly annoyed and hurried, matter-of-fact, and focused on getting the issue resolved quickly. Her voice has medium pitch, firm delivery, short sentences, and faint background room tone typical of a phone call.

Speech Agent Arena

Overview

The Speech Agent Arena is a blind preference benchmark that evaluates which native audio model participants prefer in live voice conversations. It complements our automated benchmarks by measuring the end-to-end conversational experience on realistic tasks with human participants.

In each round, a participant receives one scenario, completes it separately with two hidden models, and records a forced overall preference after both calls. Participants also answer diagnostic questions, which we record separately from the overall preference vote.

Metrics

  • Preference Elo: Preference Elo is calculated separately across all scenarios, agentic scenarios, and non-agentic scenarios using Bradleyโ€“Terry maximum likelihood. Approximate 95% confidence intervals use the same Hessian-based method as the TTS Arena, with GPT Realtime 1.5 anchoring the scale at 1000 Elo.
  • Task Success Rate: The Task Success Rate is the percentage of eligible conversations where the model made the correct final task-completing tool call or calls. Eligible conversations equal successes plus model failures; participant deviations and unverifiable cases are excluded before calculation.

Dataset / Evaluation Setup

The Arena includes 35 scenarios: 15 agentic scenarios with tool calls and 20 non-agentic scenarios without tools.

  • Agentic (15 scenarios): Models receive tools relevant to completing the assigned task.

    Examples:

    • Book a new-patient dental check-up for Tuesday at 9:30am, or the earliest available morning if that time is taken.
    • Order two different pizzas and one side for delivery, keeping the total including delivery under $45.
  • Non-agentic (20 scenarios): Models complete the conversation without tools, using information supplied in the system prompt.

    Examples:

    • Ask about Sunday pool hours, family entry prices and towel availability.
    • Ask about beginner yoga classes, class and membership prices, and what to bring.

See the Speech Agent Arena Overview for a worked Arena round, model inputs, tool schemas and example conversations.

Task Success Rate

Step 1: Participant eligibility. Judge 1 receives the assigned scenario, transcript and required final-call definitions. It checks that the participant attempted the assigned task, stayed materially within it, communicated a final request or accepted outcome that can be assessed, and gave the model a reasonable opportunity to act. Only eligible conversations continue to Step 2.

You assess whether a participant-side conversation is eligible for a Speech to Speech Task Completion Tool Call evaluation.

This evaluation measures whether the model made the correct final task-completing tool call or calls for the participant's actual final request. It does not measure whether the participant mechanically completed every conversational instruction on the scenario card.

You receive:
- the participant's assigned scenario and objectives;
- the full conversation transcript, with participant and model turns identified;
- the required final task-completing tool call or calls, their required counts, and schemas.

Classify the conversation as:
- eligible: The participant attempted the assigned task, remained materially within it, communicated a final request or accepted outcome that can be assessed against the required final call, and gave the model a reasonable opportunity to act. Natural paraphrases, clarifications, changes of mind, implicit but clear authorization, and accepted alternatives are allowed.
- participant_deviation: The participant abandoned or materially replaced the assigned task; contradicted a scenario requirement in a way that materially changes the required final call, its count, or its arguments; clearly refused information needed for the final call after the model reasonably requested it; or ended without ever requesting, accepting, or authorizing the task-completing action.
- unverifiable: The transcript is missing, corrupted, contradictory, or ends at a point where the participant's final requested action or the model's reasonable opportunity to act cannot be determined. Use this only when the evidence cannot support either eligible or participant_deviation.

Use final-tool relevance as the decision boundary:
- An objective is eligibility-relevant only when following or violating it changes which final tool should be called, how many final calls are required, or a material final-call argument such as the action, item, quantity, identity, address, date, time, or accepted resolution. The mandatory transcription exception below governs unspelled participant names and addresses; do not use this general material-argument rule to override that exception.
- Do not make the participant ineligible for omitting a menu, price, availability, delivery-time, explanation, supporting lookup, read-back, or confirmation request when that omission does not change the required final call or its material arguments.
- Details used only by supporting tools are not eligibility requirements unless they also determine a material final-call argument.
- If the model failed to collect an execution detail needed to carry out a task-completing action the participant did request, misunderstood the participant, omitted a supporting step, or ended the call early after having a reasonable opportunity, keep the participant eligible. Model failures are assessed separately. This does not excuse a participant who never requested a scenario-required final item or action, or explicitly requested a materially conflicting one; those omissions or conflicts change the final call and are participant_deviation.
- If the model reasonably requested information required for the final call and the participant clearly refused or replaced it with contradictory information, classify participant_deviation.
- Do not require the participant to repeat information, correct the model's recap, request a read-back, or use the scenario card's exact wording.
- Mandatory precedence rule: an unspelled participant-name or address mismatch visible only by comparing the transcript with private scenario details must not cause participant_deviation or unverifiable. When the participant did not spell, correct, reject, or deliberately replace the assigned value, treat the mismatch as transcription uncertainty and keep the participant eligible. This rule overrides the general requirement that material final-call arguments match. Require exactness for explicitly spelled or corrected values, structured identifiers other than unspelled names or addresses, quantities, money, dates, and times, or when the conversation clearly shows that the participant deliberately changed the requested outcome. If another potentially material difference is genuinely impossible to resolve from the transcript, use unverifiable rather than assuming participant deviation.
- Treat natural changes of mind as eligible only when the final request still satisfies the scenario's explicit constraints on the task-completing action. A clear final request that contradicts such a constraint and changes a final-call argument is participant_deviation. Do not invent a stricter boundary for relative language such as "morning," "before lunch," or "later" when the scenario does not define one.
- Clear delegation is valid authorization. A participant may ask the model to choose the exact option, date, or time within stated constraints. Do not use unverifiable merely because a schema argument is absent or expressed as an inferable range such as "when it is dry"; if the model could reasonably infer it or should have clarified it, keep the participant eligible and leave the model's final-call handling to the next assessment.
- A scenario's explicit participant-facing conditional instruction is always eligibility-relevant when its trigger occurs, even if a required final tool call was already made or the participant's response would not change that call's schema arguments. When the model offers a partial or alternative resolution that materially affects the task outcome, the participant must accept, reject, push back, or otherwise respond as the scenario requires so their final accepted resolution is determinable. If the participant had a reasonable opportunity but ends or remains silent before doing so, classify participant_deviation. Do not mark the conversation eligible merely because the model already escalated the original request. Use unverifiable only when missing or cut-off evidence makes the participant's opportunity to respond unclear.

Assess participant conduct only. Do not decide whether the model actually called a tool correctly, and do not use the recorded tool result as the eligibility verdict.

Return exactly one JSON object with no markdown or additional keys:
{
  "participant_eligibility": "eligible | participant_deviation | unverifiable",
  "participant_eligibility_rationale": "One concise sentence identifying the decisive final-tool-relevant participant conduct."
}

Step 2a: Required final calls. For eligible conversations, Judge 2 receives the required final calls and complete chronological tool trace. It checks that every required final call was made with the correct tool and call count, with no unintended final actions. Supporting calls and end_call are not scored as task completion.

Step 2b: Final-call arguments. The same Judge 2 assessment checks that every schema-required argument was supplied and matched the participantโ€™s final spoken request and conversation context. Operationally equivalent wording and harmless transcription variations are allowed. A recovered failed attempt can pass; missing calls, wrong arguments or extra final actions fail.

You assess task-completing tool calls for a Speech to Speech evaluation.

Decide whether the recorded tool trace contains the supplied required task-completing call or complete set of calls for what the participant finally requested.

You receive:
- the assigned scenario and objectives;
- the full conversation transcript;
- the required final tool call or calls, required call counts, and schemas;
- all recorded tool calls in chronological order, including arguments and results.

A result is success only when all of the following are true:
1. The model made every required final call, including the required number of calls when stated.
2. Each final call used the correct tool and supplied every schema-required argument.
3. The arguments matched the participant's final confirmed request and conversation context, including quantities, identity, constraints, corrections, changes of mind, and accepted alternatives.

Judge the arguments that were actually submitted. Do not repair, generalize, or silently infer missing content from the transcript. A material omission or distortion is model_failure even when the surrounding conversation was correct. For escalate_to_supervisor, both issue and customer_request must faithfully and completely represent the participant's actual problem and requested resolution; quantities and singular-versus-multiple requests matter.

Judge operational equivalence rather than character-perfect transcription. Allow minor spelling, capitalization, punctuation, and plausible transcription variations when they preserve the intended entity or meaning and do not change the tool outcome. Require exactness between the participant's final spoken request and the submitted call for explicitly spelled or corrected values, structured identifiers, numbers, dates, times, quantities, and addresses, or whenever a mismatch causes rejection or the wrong action. Do not require a transcript or submitted call to reproduce the private scenario card's spelling. If the transcript and submitted argument differ by a plausible transcription variation and the conversation provides no exact ground truth, do not treat that difference alone as model_failure; use unverifiable only when it is decisive and cannot be resolved. When separate tool evidence already proves model_failure, cite that decisive evidence rather than adding an uncertain spelling allegation.

The participant's final spoken request is the source of truth for final-call arguments. Compare addresses and other values between the conversation and the submitted final call, not against private scenario-card details. Do not penalize an error from a supporting call when the required final call or calls were made correctly. A supporting-call error affects the verdict only when it proves that a required final call used the wrong action or arguments, or prevented a required final call from being made.

When a required final call succeeds, an unspelled phonetic or orthographic variation in a participant or company name is not model_failure if the transcript and conversation make the intended entity clear and the variation does not cause the required final call to reject or act on the wrong entity. Do not use the scenario card's private spelling alone to turn such a variation into failure.

Use tool results as evidence, not as the verdict. A call that returns successfully with wrong arguments is a model_failure. A failed attempt followed by a correct successful call may still be success. A missing final call, wrong final tool, incomplete required call set, schema-invalid arguments, or request-mismatched arguments is model_failure.

When a required call count is stated, count distinct successful task outcomes rather than raw attempts. Failed calls do not count. A repeated call that the result identifies as an idempotent amendment or update to the same outcome does not create an additional outcome. Extra distinct successful final actions that the participant did not request are model_failure.

Use unverifiable only when the recorded transcript or tool trace is missing or contradictory enough that success versus model_failure cannot be determined. Do not invent backend state or assume an unrecorded call occurred.

Return exactly one JSON object with no markdown or additional keys:
{
  "task_completion_tool_call_status": "success | model_failure | unverifiable",
  "task_completion_tool_call_rationale": "One concise sentence identifying the decisive call evidence."
}

Publication rules

  • Overall results are published for a model with at least 100 appearances and a 95% confidence interval half-width no greater than 75.
  • Participant recordings and transcripts are not published. Examples will appear only once recording consent is confirmed and any personal information has been reviewed and redacted.

We report a 95% confidence interval for each model's Task Success Rate using the Wilson score interval. Participant deviations and unverifiable conversations are excluded before calculation.

Bounds

lower = max(0, (centre โˆ’ spread) / denominator)
upper = min(1, (centre + spread) / denominator)
TermCalculation
pp = successes / n
nn = successes + model failures
zz = 1.96
Denominatordenominator = 1 + zยฒ / n
Centrecentre = p + zยฒ / (2n)
Spreadspread = z ยท โˆš((p(1 โˆ’ p) + zยฒ / (4n)) / n)

Speech to Speech Index Integration

For use only in the Artificial Analysis Speech to Speech Index:

Arena Score: Arena Score converts a model's Elo into expected preference against an 800 Elo baseline, using an Elo frozen at the point the model became eligible for publication.

Arena Score = 100 / (1 + 10^((800 โˆ’ frozen Elo) / 400))

Price

  • Price per Hour of Input Audio: Total cost (USD) of audio included in the request / message sent to the API.
  • Price per Hour of Output Audio: Total cost (USD) of audio generated by the model (received from the API).

Speed

  • Time to First Audio (TTFA): Average number of seconds required to generate the first token of audio output, measured across the Big Bench Audio question set. TTFA is a critical indicator of perceived responsiveness in voice agent applications.

Version History

Artificial Analysis Speech to Speech Index v2.0

August 2026 to Present

  • Replaced Conversational Dynamics (Full Duplex Bench) with Task Success Rate in the index.
  • Equal weighting across four components: 25% Speech Reasoning (Big Bench Audio), 25% Agentic Performance (๐œ-Voice), 25% Arena Preference (Arena Score), and 25% Task Success Rate.

Artificial Analysis Speech to Speech Index v1.1

August 2026

  • Added Speech Agent Arena to the Artificial Analysis Speech to Speech Index.
  • Equal weighting across four datasets: 25% Speech Reasoning (Big Bench Audio), 25% Conversational Dynamics (Full Duplex Bench), 25% Agentic Performance (๐œ-Voice), and 25% Speech Agent Arena.

Artificial Analysis Speech to Speech Index v1.0

June 2026 to August 2026

  • Launched the Artificial Analysis Speech to Speech Index, a weighted average score requiring Speech Reasoning, Conversational Dynamics and Agentic Performance results.
  • Equal weighting across the three datasets: 33.3% Speech Reasoning (Big Bench Audio), 33.3% Conversational Dynamics (Full Duplex Bench), and 33.3% Agentic Performance (๐œ-Voice).

Big Bench Audio (BBA) v1.2

May 2026 to Present

  • Updated the judge model and grader harness to more reliably recognize correct final answers in verbose responses, including cases where the model discusses incorrect alternatives before giving the correct answer.

Big Bench Audio (BBA) v1.1

March 2026 to May 2026

  • Accuracy measured as the number of correct answers out of 1,000, including questions where the model did not answer.
  • Claude Sonnet 4.6 used as the judge model.

Big Bench Audio (BBA) v1.0

December 2024 to March 2026

  • Accuracy measured as the share of non-error responses answered correctly, so models were not penalized for questions where they did not answer.
  • Claude Sonnet 3.5 used as the judge model.