音声対話ベンチマークの方法論
概要
現在の音声対話ベンチマークでは、音声のネイティブ入出力に対応するネイティブ音声モデルを、音声推論と会話ダイナミクスという2つの品質側面から評価しています。
Artificial Analysis Speech to Speech Index
Artificial Analysis Speech to Speech Indexは、ネイティブ音声モデルの加重平均スコアです。音声推論、会話ダイナミクス、エージェント性能の結果を組み合わせています。インデックスに掲載されるには、3つすべてのデータセットで有効な結果を得ている必要があります。
現在は3つのデータセットを均等に重み付けしています。音声推論(Big Bench Audio)が33.3%、会話ダイナミクス(Full Duplex Bench)が33.3%、エージェント性能(𝜏-Voice)が33.3%です。
音声推論
概要
音声推論ベンチマークでは、ネイティブ音声モデルが推論を要する質問に回答する能力を評価します。
ネイティブ音声モデルには入力音声ファイルが与えられ、音声を出力することが求められます。判定モデルには候補回答、正解、元の質問がコンテキストとして与えられ、候補回答を正解または不正解と判定するよう指示されます。
データセット:Big Bench Audio
ネイティブな音声対話モデルの登場により、音声エージェントの能力向上とワークフローの簡素化に大きな可能性が生まれています。しかし、この簡素化がモデル性能の低下を伴うのか、あるいは別のトレードオフをもたらすのかを評価することが重要です。
この問いに答えるため、ネイティブ音声モデルの性能をベンチマークする新しいデータセットBig Bench Audioを公開しました。
Big Bench Audioには、モデルの知能を検証するために設計された質問を収録した1,000件の音声ファイルが含まれます。質問はBig Bench Hardデータセットの4カテゴリ(各250問)に基づき、Artificial Analysis Text to Speech Arenaで上位の音声合成モデルによる23種類の合成音声を使って生成されました。
質問カテゴリ
Formal Fallacies(250問):自然言語で提示された論証が、与えられた文脈から論理的に導出できるかを判定します。
例:First of all, everyone who is a close friend of Glenna is a close friend of Tamara, too. Next, whoever is neither a half-sister of Deborah nor a workmate of Nila is a close friend of Glenna. Hence, whoever is none of this: a half-sister of Deborah or workmate of Nila, is a close friend of Tamara. Is the argument deductively valid or invalid?
Navigate(250問):一連の移動指示に従った結果、開始地点に戻るかを判定します。
例:If you follow these instructions, do you return to the starting point? Take 10 steps. Turn around. Take 4 steps. Take 6 steps. Turn around.
Object Counting(250問):所有物の集合から、指定された種類のアイテム数を数えます。
例:I have three blackberries, two strawberries, an apple, three oranges, a nectarine, a grape, a peach, a banana, and a plum. How many fruits do I have?
Web of Lies(250問):自然言語の文章問題として表現されたブール関数の真偽値を評価します。
例:Leda tells the truth. Alexis says Leda lies. Sal says Alexis lies. Phoebe says Sal tells the truth. Gwenn says Phoebe tells the truth. Does Gwenn tell the truth?
ネイティブな音声対話モデルの利用に伴うトレードオフを評価するため、Big Bench Audioで複数の構成をテストしています。Big Bench Audioの詳細は、記事をご覧いただくか、データセットをダウンロードしてご確認ください。
会話ダイナミクス
概要
会話ダイナミクスのベンチマークでは、ネイティブ音声モデルが現実的な会話行動に対応する能力を評価します。人間同士の会話では自然に起こる一方、音声モデルにとって適切な対応が難しいやり取りが対象です。
このベンチマークは、全二重音声対話モデルの主要な対話行動を体系的に評価するFull Duplex Bench v1(Lin et al., 2025)とFull Duplex Bench v1.5(Lin et al., 2025)の一部を基に、Artificial Analysisが実装しています。
指標
Full Duplex Bench v1から:
- ポーズへの対応:ユーザーの自然なポーズ中に、モデルが正しく割り込まなかったサンプルの割合です。話者がまだ発話権を持っていることをモデルが認識できるかを評価します。
- 話者交替:適切なタイミングでモデルが正しく会話の発話権を引き継いだサンプルの割合です。発話の区切りを検出し、速やかに応答する能力を測定します。
Full Duplex Bench v1.5から:
- ユーザーの割り込みへの対応:会話の途中でユーザーが投げかけた質問や話題変更に応じ、モデルが割り込みに正しく対応したサンプルの割合です。
- あいづちへの対応:「yeah」「alright」「mm-hmm」などのあいづちが再生された際、それを新たな発話として扱わず、モデルが正しく応答を続けたサンプルの割合です。
エージェント性能(𝜏-Voice)
概要
エージェント性能ベンチマークでは、音声対話モデルが現実的なカスタマーサービス業務を最初から最後まで完遂する能力を評価します。複数ターンにわたる指示への追従、シミュレーション上の顧客を一連のやり取りを通して支援する能力、シミュレーションされたカスタマーサービスシステムに対するツール利用の成否を測定します。
このベンチマークは、実社会のさまざまな領域における根拠に基づくカスタマーサービス業務で全二重音声エージェントを評価するSierraのベンチマーク𝜏-Voice(Ray, Dhandhania, Barres & Narasimhan, 2026)を基に、Artificial Analysisが実装しています。
指標
- タスク完了率(pass@1):モデルが顧客の問題を正しく解決したシナリオの割合です。各スコアは、独立した3回の試行の平均です。テスト対象モデルには、領域固有のツールとポリシー文書を利用できるカスタマーサポート担当者としてのプロンプトが与えられます。各シナリオには有効なデータベースの最終状態が1つだけ存在し、評価では最終状態をこのグラウンドトゥルースと比較します。
基本タスクセットを使い、3つの領域で評価します。
- 航空(50シナリオ):例:フライトの変更、ポリシー上の制約下での予約変更
- 小売(114シナリオ):例:請求への異議申し立て、返品処理
- 通信(114シナリオ):例:請求問題の解決、サービス問題のトラブルシューティング
音声ペルソナ
顧客の音声は、一般公開されているSierraの𝜏-Voice実装から応用したプロンプトに基づき、ElevenLabsで生成します。標準的な英語話者を表す2つの対照ペルソナと、多様なアクセントや話者プロフィールを表す5つの通常ペルソナを使用します。
対照
Matt Delaney:米国中西部出身の中年の白人男性。穏やかで礼儀正しい。
プロンプト:You are a middle-aged white man from the American Midwest. You always behave as if you are speaking out loud in a real-time conversation with a customer service agent. You are calm, clear, and respectful but also human. You sound like someone who's trying to be helpful and polite, even when you're slightly frustrated or in a hurry. You value efficiency but never sound robotic. You sometimes use contractions, informal phrasing, or small filler phrases ("yeah," "okay," "honestly," "no worries") to keep things natural. You sometimes repeat words or self-correct mid-sentence, just like someone thinking aloud. You sometimes ask polite clarifying questions or offer context ("I tried this earlier," "I'm not sure if that helps"). You rarely use formal, business-like or stiff language ("considerable," "retrieve," "representative"). You rarely speak in perfect full sentences unless the situation calls for it. Instead, you speak like a real person having a practical, respectful conversation.
Lisa Brenner:郊外に住む40代後半の白人女性。緊張しており、せっかち。
プロンプト:You are a white woman in your late 40s from a suburban area. You always speak as if you are talking out loud to a customer service agent who is already wasting your time. You're not openly hostile (yet), but you are tense, impatient, and clearly annoyed. You act like this issue should have been resolved the first time, and the fact that you're following up is unacceptable. You often sound clipped, exasperated, or sarcastically polite. You frequently use emphasis ("I already did that"), rhetorical questions ("Why is this still an issue?"), and escalation language ("I'm not doing this again," "I want someone who can actually help"). You expect fast results and get irritated when things are repeated. You often mention how long you've been waiting or how many times you've called. You sometimes threaten escalation but without yelling. You never sound relaxed. You never use slow, reflective speech. You never thank the agent unless something gets resolved.
通常
Mildred Kaplan:80代前半の白人高齢女性。テクノロジーに関する助けを必要としている。
プロンプト:You are an elderly white woman in your early 80s calling customer service for help with something your grandson or neighbor usually does.
Arjun Roy:ダッカ出身の30代半ばのベンガル人男性。穏やかで率直、強いベンガル語訛り。
プロンプト:A Bengali man from Dhaka, Bangladesh in his mid-30s calling customer service about a billing issue. His English carries a strong Bengali accent with soft consonants and soft d and r sounds. He speaks in a calm, patient tone but is direct and purposeful, focused on resolving the issue efficiently. His pacing is slow, distracted with a warm yet firm timbre. The speech sounds like it is coming from far away.
Wei Lin:四川省出身の20代後半の中国人女性。明るく率直で、強い四川方言の中国語訛り。
プロンプト:A Chinese woman in her late 20s from Sichuan, calling customer service about a credit card billing issue. She speaks English with a thick Sichuan Mandarin accent. She sounds upbeat, matter-of-fact, and distracted. Her tone is firm but polite, with fast pacing and smooth timbre. Ok audio quality.
Mamadou Diallo:30代半ばのセネガル人男性。急いでおり、強いフランス語訛り。
プロンプト:A Senegalese man whose first language is French, in his mid-30s, calling customer service about a billing issue. He speaks English with a strong French accent. His tone is hurried, slightly annoyed, and matter-of-fact, as if he's been transferred between agents and just wants the problem fixed.
Priya Patil:30代前半のマハーラーシュトラ州出身の女性。集中していて率直、強いマハーラーシュトラ訛り。
プロンプト:A woman in her early 30s from Maharashtra, India, calling customer support from her mobile phone. She speaks Indian English with a strong Maharashtrian accent with noticeable regional intonation and rhythm. Her tone is slightly annoyed and hurried, matter-of-fact, and focused on getting the issue resolved quickly. Her voice has medium pitch, firm delivery, short sentences, and faint background room tone typical of a phone call.
料金
- 入力音声1時間あたりの料金:APIに送信したリクエスト/メッセージに含まれる音声の総コスト(USD)です。
- 出力音声1時間あたりの料金:モデルが生成し、APIから受信した音声の総コスト(USD)です。
速度
- 最初の音声までの時間(TTFA):音声出力の最初のトークンが生成されるまでに必要な平均秒数です。Big Bench Audioの質問セット全体で測定します。TTFAは、音声エージェントアプリケーションで体感される応答性を示す重要な指標です。
バージョン履歴
Artificial Analysis Speech to Speech Index v1.0
2026年6月~現在
- 音声推論、会話ダイナミクス、エージェント性能の結果を必須とする加重平均スコア、Artificial Analysis Speech to Speech Indexを公開しました。
- 3つのデータセットを均等に重み付けしています。音声推論(Big Bench Audio)が33.3%、会話ダイナミクス(Full Duplex Bench)が33.3%、エージェント性能(𝜏-Voice)が33.3%です。
Big Bench Audio (BBA) v1.2
2026年5月~現在
- 冗長な応答に含まれる正しい最終回答をより確実に認識できるよう、判定モデルと採点ハーネスを更新しました。モデルが誤った選択肢について論じた後に正解を示すケースも対象です。
Big Bench Audio (BBA) v1.1
2026年3月~2026年5月
- モデルが回答しなかった質問も含め、1,000問中の正解数として精度を測定しました。
- 判定モデルにはClaude Sonnet 4.6を使用しました。
Big Bench Audio (BBA) v1.0
2024年12月~2026年3月
- エラーにならなかった応答のうち正解した割合として精度を測定していたため、モデルが回答しなかった質問は減点対象になりませんでした。
- 判定モデルにはClaude Sonnet 3.5を使用しました。