音声対話ベンチマークの方法論
概要
現在の音声対話ベンチマークでは、音声のネイティブ入出力に対応するネイティブ音声モデルを、音声推論と会話ダイナミクスという2つの品質側面から評価しています。
Artificial Analysis Speech to Speech Index
Artificial Analysis Speech to Speech Indexは、ネイティブ音声モデルの均等加重スコアです。音声推論、エージェント性能、Arena Preference、Task Success Rateの結果を組み合わせています。インデックスに掲載されるには、4つすべてのコンポーネントで有効な結果を得ている必要があります。
現在の重み付けは4つのコンポーネントで均等です:音声推論(Big Bench Audio)25%、エージェント性能(𝜏-Voice)25%、Arena Preference(Arena Score)25%、Task Success Rate 25%。
音声推論
概要
音声推論ベンチマークでは、ネイティブ音声モデルが推論を要する質問に回答する能力を評価します。
ネイティブ音声モデルには入力音声ファイルが与えられ、音声を出力することが求められます。判定モデルには候補回答、正解、元の質問がコンテキストとして与えられ、候補回答を正解または不正解と判定するよう指示されます。
データセット:Big Bench Audio
ネイティブな音声対話モデルの登場により、音声エージェントの能力向上とワークフローの簡素化に大きな可能性が生まれています。しかし、この簡素化がモデル性能の低下を伴うのか、あるいは別のトレードオフをもたらすのかを評価することが重要です。
この問いに答えるため、ネイティブ音声モデルの性能をベンチマークする新しいデータセットBig Bench Audioを公開しました。
Big Bench Audioには、モデルの知能を検証するために設計された質問を収録した1,000件の音声ファイルが含まれます。質問はBig Bench Hardデータセットの4カテゴリ(各250問)に基づき、Artificial Analysis Text to Speech Arenaで上位の音声合成モデルによる23種類の合成音声を使って生成されました。
質問カテゴリ
Formal Fallacies(250問):自然言語で提示された論証が、与えられた文脈から論理的に導出できるかを判定します。
例:First of all, everyone who is a close friend of Glenna is a close friend of Tamara, too. Next, whoever is neither a half-sister of Deborah nor a workmate of Nila is a close friend of Glenna. Hence, whoever is none of this: a half-sister of Deborah or workmate of Nila, is a close friend of Tamara. Is the argument deductively valid or invalid?
Navigate(250問):一連の移動指示に従った結果、開始地点に戻るかを判定します。
例:If you follow these instructions, do you return to the starting point? Take 10 steps. Turn around. Take 4 steps. Take 6 steps. Turn around.
Object Counting(250問):所有物の集合から、指定された種類のアイテム数を数えます。
例:I have three blackberries, two strawberries, an apple, three oranges, a nectarine, a grape, a peach, a banana, and a plum. How many fruits do I have?
Web of Lies(250問):自然言語の文章問題として表現されたブール関数の真偽値を評価します。
例:Leda tells the truth. Alexis says Leda lies. Sal says Alexis lies. Phoebe says Sal tells the truth. Gwenn says Phoebe tells the truth. Does Gwenn tell the truth?
ネイティブな音声対話モデルの利用に伴うトレードオフを評価するため、Big Bench Audioで複数の構成をテストしています。Big Bench Audioの詳細は、記事をご覧いただくか、データセットをダウンロードしてご確認ください。
会話ダイナミクス
概要
会話ダイナミクスのベンチマークでは、ネイティブ音声モデルが現実的な会話行動に対応する能力を評価します。人間同士の会話では自然に起こる一方、音声モデルにとって適切な対応が難しいやり取りが対象です。
このベンチマークは、全二重音声対話モデルの主要な対話行動を体系的に評価するFull Duplex Bench v1(Lin et al., 2025)とFull Duplex Bench v1.5(Lin et al., 2025)の一部を基に、Artificial Analysisが実装しています。
指標
Full Duplex Bench v1から:
- ポーズへの対応:ユーザーの自然なポーズ中に、モデルが正しく割り込まなかったサンプルの割合です。話者がまだ発話権を持っていることをモデルが認識できるかを評価します。
- 話者交替:適切なタイミングでモデルが正しく会話の発話権を引き継いだサンプルの割合です。発話の区切りを検出し、速やかに応答する能力を測定します。
Full Duplex Bench v1.5から:
- ユーザーの割り込みへの対応:会話の途中でユーザーが投げかけた質問や話題変更に応じ、モデルが割り込みに正しく対応したサンプルの割合です。
- あいづちへの対応:「yeah」「alright」「mm-hmm」などのあいづちが再生された際、それを新たな発話として扱わず、モデルが正しく応答を続けたサンプルの割合です。
エージェント性能(𝜏-Voice)
概要
エージェント性能ベンチマークでは、音声対話モデルが現実的なカスタマーサービス業務を最初から最後まで完遂する能力を評価します。複数ターンにわたる指示への追従、シミュレーション上の顧客を一連のやり取りを通して支援する能力、シミュレーションされたカスタマーサービスシステムに対するツール利用の成否を測定します。
このベンチマークは、実社会のさまざまな領域における根拠に基づくカスタマーサービス業務で全二重音声エージェントを評価するSierraのベンチマーク𝜏-Voice(Ray, Dhandhania, Barres & Narasimhan, 2026)を基に、Artificial Analysisが実装しています。
指標
- タスク完了率(pass@1):モデルが顧客の問題を正しく解決したシナリオの割合です。各スコアは、独立した3回の試行の平均です。テスト対象モデルには、領域固有のツールとポリシー文書を利用できるカスタマーサポート担当者としてのプロンプトが与えられます。各シナリオには有効なデータベースの最終状態が1つだけ存在し、評価では最終状態をこのグラウンドトゥルースと比較します。
基本タスクセットを使い、3つの領域で評価します。
- 航空(50シナリオ):例:フライトの変更、ポリシー上の制約下での予約変更
- 小売(114シナリオ):例:請求への異議申し立て、返品処理
- 通信(114シナリオ):例:請求問題の解決、サービス問題のトラブルシューティング
音声ペルソナ
顧客の音声は、一般公開されているSierraの𝜏-Voice実装から応用したプロンプトに基づき、ElevenLabsで生成します。標準的な英語話者を表す2つの対照ペルソナと、多様なアクセントや話者プロフィールを表す5つの通常ペルソナを使用します。
対照
Matt Delaney:米国中西部出身の中年の白人男性。穏やかで礼儀正しい。
プロンプト:You are a middle-aged white man from the American Midwest. You always behave as if you are speaking out loud in a real-time conversation with a customer service agent. You are calm, clear, and respectful but also human. You sound like someone who's trying to be helpful and polite, even when you're slightly frustrated or in a hurry. You value efficiency but never sound robotic. You sometimes use contractions, informal phrasing, or small filler phrases ("yeah," "okay," "honestly," "no worries") to keep things natural. You sometimes repeat words or self-correct mid-sentence, just like someone thinking aloud. You sometimes ask polite clarifying questions or offer context ("I tried this earlier," "I'm not sure if that helps"). You rarely use formal, business-like or stiff language ("considerable," "retrieve," "representative"). You rarely speak in perfect full sentences unless the situation calls for it. Instead, you speak like a real person having a practical, respectful conversation.
Lisa Brenner:郊外に住む40代後半の白人女性。緊張しており、せっかち。
プロンプト:You are a white woman in your late 40s from a suburban area. You always speak as if you are talking out loud to a customer service agent who is already wasting your time. You're not openly hostile (yet), but you are tense, impatient, and clearly annoyed. You act like this issue should have been resolved the first time, and the fact that you're following up is unacceptable. You often sound clipped, exasperated, or sarcastically polite. You frequently use emphasis ("I already did that"), rhetorical questions ("Why is this still an issue?"), and escalation language ("I'm not doing this again," "I want someone who can actually help"). You expect fast results and get irritated when things are repeated. You often mention how long you've been waiting or how many times you've called. You sometimes threaten escalation but without yelling. You never sound relaxed. You never use slow, reflective speech. You never thank the agent unless something gets resolved.
通常
Mildred Kaplan:80代前半の白人高齢女性。テクノロジーに関する助けを必要としている。
プロンプト:You are an elderly white woman in your early 80s calling customer service for help with something your grandson or neighbor usually does.
Arjun Roy:ダッカ出身の30代半ばのベンガル人男性。穏やかで率直、強いベンガル語訛り。
プロンプト:A Bengali man from Dhaka, Bangladesh in his mid-30s calling customer service about a billing issue. His English carries a strong Bengali accent with soft consonants and soft d and r sounds. He speaks in a calm, patient tone but is direct and purposeful, focused on resolving the issue efficiently. His pacing is slow, distracted with a warm yet firm timbre. The speech sounds like it is coming from far away.
Wei Lin:四川省出身の20代後半の中国人女性。明るく率直で、強い四川方言の中国語訛り。
プロンプト:A Chinese woman in her late 20s from Sichuan, calling customer service about a credit card billing issue. She speaks English with a thick Sichuan Mandarin accent. She sounds upbeat, matter-of-fact, and distracted. Her tone is firm but polite, with fast pacing and smooth timbre. Ok audio quality.
Mamadou Diallo:30代半ばのセネガル人男性。急いでおり、強いフランス語訛り。
プロンプト:A Senegalese man whose first language is French, in his mid-30s, calling customer service about a billing issue. He speaks English with a strong French accent. His tone is hurried, slightly annoyed, and matter-of-fact, as if he's been transferred between agents and just wants the problem fixed.
Priya Patil:30代前半のマハーラーシュトラ州出身の女性。集中していて率直、強いマハーラーシュトラ訛り。
プロンプト:A woman in her early 30s from Maharashtra, India, calling customer support from her mobile phone. She speaks Indian English with a strong Maharashtrian accent with noticeable regional intonation and rhythm. Her tone is slightly annoyed and hurried, matter-of-fact, and focused on getting the issue resolved quickly. Her voice has medium pitch, firm delivery, short sentences, and faint background room tone typical of a phone call.
Speech Agent Arena
概要
Speech Agent Arenaはブラインド選好ベンチマークです。各ラウンドで参加者は1つのシナリオを受け取り、非公開の2モデルとライブ音声通話で会話したうえで、全体選好を必須で記録し、5つの診断的比較を行います。引き分けは診断質問に限って認められます。
全体、エージェント、非エージェント、診断の各リーダーボードは、Bradley-Terry最尤推定により独立に推定され、Bradley-Terry推定のヘシアン対角に基づく95%信頼区間を伴います。GPT Realtime 1.5がスケールを1000 Eloに固定します。
公開基準
- 全体結果は、出現回数が100以上かつ95%信頼区間の半幅が75以下のモデルについて公開されます。
- 診断結果は、出現回数が50以上で、かつ比較グラフが連結している場合に限り公開されます。
- 参加者の録音と文字起こしは公開されません。例は、録音の同意が確認され、個人情報の確認と修正が完了した後にのみ掲載されます。
Task Success Rate
15 のエージェント型シナリオには、モデルが各タスクを完了するために実行する必要があるツール呼び出しが含まれます。別の評価パイプラインが、各会話でモデルが正しい最終タスク完了ツール呼び出しを行ったかどうかを評価します。
評価は 2 段階のプロセスで行います。ステップ 1 では参加者の適格性を確認します:LLM ジャッジが、参加者が割り当てタスクに取り組み、実質的にその範囲内にとどまり、モデルに行動する妥当な機会を与えたかを確認します。適格な会話のみがステップ 2 に進みます。ステップ 2 ではツール呼び出しの正確性を確認します:2 番目の LLM ジャッジが、必要な最終呼び出しがすべて正しいツールと回数で行われ、参加者の最終的な発話要求にスキーマが要求する引数が一致しているかを検証します。
Task Success Rate は、適格な会話のうちモデルが正しい最終タスク完了ツール呼び出しを行った割合です。適格な会話数 = 成功数 + モデルの失敗数。参加者の逸脱と検証不能なケースは計算前に除外されます。
各モデルの Task Success Rate に対する 95% 信頼区間は、z = 1.96 の Wilson スコア区間を用いて計算されます。p を観測された成功率、n を適格な会話数とすると、上下限は以下のとおりです:denom = 1 + z²/n, centre = p + z²/(2n), spread = z · √((p(1−p) + z²/(4n)) / n), lower = max(0, (centre − spread) / denom), upper = min(1, (centre + spread) / denom)。
Wilson スコア区間は、Bradley-Terry推定のヘシアン対角から得られる Elo の信頼区間とは異なります。小さなサンプルサイズや 0 または 1 に近い割合において、正規近似よりも良好なカバレッジを提供します。
Arena Score
Arena Scoreは、モデルが公開対象となった時点で固定したEloを用いて、800 Eloのベースラインに対する期待選好度に変換したものです。
診断Eloは別途報告され、全体選好結果やArena Scoreには反映されません。
Arena Score = 100 / (1 + 10^((800 − frozen Elo) / 400))
料金
- 入力音声1時間あたりの料金:APIに送信したリクエスト/メッセージに含まれる音声の総コスト(USD)です。
- 出力音声1時間あたりの料金:モデルが生成し、APIから受信した音声の総コスト(USD)です。
速度
- 最初の音声までの時間(TTFA):音声出力の最初のトークンが生成されるまでに必要な平均秒数です。Big Bench Audioの質問セット全体で測定します。TTFAは、音声エージェントアプリケーションで体感される応答性を示す重要な指標です。
バージョン履歴
Artificial Analysis Speech to Speech Index v2.0
2026年8月~現在
- インデックスにおいて会話ダイナミクス(Full Duplex Bench)をTask Success Rateに置き換えました。
- 4つのコンポーネントを均等に重み付け:音声推論(Big Bench Audio)25%、エージェント性能(𝜏-Voice)25%、Arena Preference(Arena Score)25%、Task Success Rate 25%。
Artificial Analysis Speech to Speech Index v1.1
2026年8月
- Artificial Analysis Speech to Speech IndexにSpeech Agent Arenaを追加しました。
- 4つのデータセットを均等に重み付け:音声推論(Big Bench Audio)25%、会話ダイナミクス(Full Duplex Bench)25%、エージェント性能(𝜏-Voice)25%、Speech Agent Arena 25%。
Artificial Analysis Speech to Speech Index v1.0
2026年6月~現在
- 音声推論、会話ダイナミクス、エージェント性能の結果を必須とする加重平均スコア、Artificial Analysis Speech to Speech Indexを公開しました。
- 3つのデータセットを均等に重み付けしています。音声推論(Big Bench Audio)が33.3%、会話ダイナミクス(Full Duplex Bench)が33.3%、エージェント性能(𝜏-Voice)が33.3%です。
Big Bench Audio (BBA) v1.2
2026年5月~現在
- 冗長な応答に含まれる正しい最終回答をより確実に認識できるよう、判定モデルと採点ハーネスを更新しました。モデルが誤った選択肢について論じた後に正解を示すケースも対象です。
Big Bench Audio (BBA) v1.1
2026年3月~2026年5月
- モデルが回答しなかった質問も含め、1,000問中の正解数として精度を測定しました。
- 判定モデルにはClaude Sonnet 4.6を使用しました。
Big Bench Audio (BBA) v1.0
2024年12月~2026年3月
- エラーにならなかった応答のうち正解した割合として精度を測定していたため、モデルが回答しなかった質問は減点対象になりませんでした。
- 判定モデルにはClaude Sonnet 3.5を使用しました。