음성 대 음성 벤치마킹 방법론
개요
현재 음성 대 음성 벤치마크에서는 네이티브 오디오 입력과 출력을 지원하는 네이티브 오디오 모델을 음성 추론과 대화 역학이라는 두 가지 품질 차원에서 평가합니다.
Artificial Analysis Speech to Speech Index
Artificial Analysis Speech to Speech Index는 네이티브 오디오 모델의 가중 평균 점수입니다. 음성 추론, 대화 역학, 에이전트 성능 결과를 종합합니다. 세 데이터 세트 모두에서 유효한 결과가 있는 모델만 이 지수에 포함됩니다.
현재 세 데이터 세트에는 동일한 가중치를 적용합니다. 음성 추론(Big Bench Audio) 33.3%, 대화 역학(Full Duplex Bench) 33.3%, 에이전트 성능(𝜏-Voice) 33.3%입니다.
음성 추론
개요
음성 추론 벤치마크는 네이티브 오디오 모델이 추론형 질문에 답하는 능력을 평가합니다.
네이티브 오디오 모델에는 입력 오디오 파일이 제공되며, 모델은 출력 오디오를 생성해야 합니다. 심사 모델에는 후보 답변, 공식 답변, 원래 질문이 문맥으로 제공되며 후보 답변이 정답인지 오답인지 판정하도록 지시합니다.
데이터 세트: Big Bench Audio
네이티브 오디오 대 오디오 모델의 등장은 음성 에이전트의 역량을 높이고 워크플로를 간소화할 흥미로운 기회를 제공합니다. 하지만 이러한 간소화로 인해 모델 성능이 저하되거나 다른 상충 관계가 발생하는지 평가하는 것이 중요합니다.
이 질문에 답하고자 네이티브 오디오 모델의 성능을 벤치마킹하는 새 데이터 세트인 Big Bench Audio를 공개했습니다.
Big Bench Audio에는 모델의 지능을 시험하기 위해 설계된 질문을 담은 오디오 파일 1,000개가 있습니다. 질문은 Big Bench Hard 데이터 세트의 네 가지 범주(각 250문항)를 기반으로 하며, Artificial Analysis Text to Speech Arena에서 상위권에 오른 텍스트 음성 변환 모델의 합성 음성 23개를 사용해 생성했습니다.
질문 범주
형식적 오류(250문항): 비형식적으로 제시된 논증이 주어진 문맥에서 논리적으로 도출될 수 있는지 판단합니다.
예시: First of all, everyone who is a close friend of Glenna is a close friend of Tamara, too. Next, whoever is neither a half-sister of Deborah nor a workmate of Nila is a close friend of Glenna. Hence, whoever is none of this: a half-sister of Deborah or workmate of Nila, is a close friend of Tamara. Is the argument deductively valid or invalid?
경로 탐색(250문항): 일련의 이동 단계를 따른 에이전트가 출발점으로 돌아오는지 판단합니다.
예시: If you follow these instructions, do you return to the starting point? Take 10 steps. Turn around. Take 4 steps. Take 6 steps. Turn around.
객체 수 세기(250문항): 소지품 모음에서 특정 항목 유형의 개수를 셉니다.
예시: I have three blackberries, two strawberries, an apple, three oranges, a nectarine, a grape, a peach, a banana, and a plum. How many fruits do I have?
거짓말의 연결망(250문항): 자연어 서술형 문제로 표현된 불리언 함수의 참과 거짓을 판단합니다.
예시: Leda tells the truth. Alexis says Leda lies. Sal says Alexis lies. Phoebe says Sal tells the truth. Gwenn says Phoebe tells the truth. Does Gwenn tell the truth?
네이티브 음성 대 음성 모델 사용에 따른 상충 관계를 평가하기 위해 Big Bench Audio에서 여러 구성을 시험합니다. Big Bench Audio에 관해 자세히 알아보려면 관련 글을 읽거나 데이터 세트를 직접 다운로드하세요.
대화 역학
개요
대화 역학 벤치마크는 네이티브 오디오 모델이 현실적인 대화 행동을 처리하는 능력을 평가합니다. 이러한 상호작용은 사람의 대화에서는 자연스럽게 일어나지만 음성 모델이 올바르게 처리하기는 어렵습니다.
이 벤치마크는 Artificial Analysis가 Full Duplex Bench v1(Lin et al., 2025)과 Full Duplex Bench v1.5(Lin et al., 2025)의 일부를 바탕으로 구현했습니다. 두 벤치마크는 전이중 음성 대화 모델의 핵심 상호작용 행동을 체계적으로 평가합니다.
지표
Full Duplex Bench v1에서 가져온 지표:
- 일시 정지 처리: 사용자가 자연스럽게 말을 멈춘 동안 모델이 올바르게 끼어들지 않은 샘플의 비율입니다. 모델이 발언권이 여전히 화자에게 있음을 인식하는지 평가합니다.
- 발언권 전환: 적절한 때에 모델이 올바르게 발언권을 넘겨받은 샘플의 비율입니다. 모델이 발화의 경계를 감지하고 신속하게 응답하는 능력을 측정합니다.
Full Duplex Bench v1.5에서 가져온 지표:
- 사용자 끼어들기 처리: 대화 도중 사용자가 제기한 질문이나 화제 전환 등 사용자의 끼어들기에 모델이 올바르게 대응한 샘플의 비율입니다.
- 백채널 처리: "yeah", "alright", "mm-hmm" 같은 백채널이 재생되었을 때 이를 새로운 발화로 간주하지 않고 모델이 응답을 올바르게 이어간 샘플의 비율입니다.
에이전트 성능(𝜏-Voice)
개요
에이전트 성능 벤치마크는 음성 대 음성 모델이 현실적인 고객 서비스 작업을 처음부터 끝까지 완료하는 능력을 평가합니다. 여러 턴에 걸친 지시 수행, 전체 상호작용에 걸쳐 시뮬레이션 고객을 지원하는 능력, 시뮬레이션 고객 서비스 시스템을 상대로 한 성공적인 도구 사용을 측정합니다.
이 벤치마크는 Sierra가 만든 𝜏-Voice(Ray, Dhandhania, Barres & Narasimhan, 2026)를 바탕으로 Artificial Analysis가 구현했습니다. 𝜏-Voice는 실제 분야의 근거 기반 고객 서비스 작업에서 전이중 음성 에이전트를 평가합니다.
지표
- 작업 완료율(pass@1): 모델이 고객의 문제를 올바르게 해결한 시나리오의 비율입니다. 각 점수는 독립적으로 수행한 세 번의 시험 평균입니다. 테스트 모델에는 분야별 도구와 정책 문서에 접근할 수 있는 고객 지원 에이전트 역할을 부여합니다. 각 시나리오에는 유효한 데이터베이스 최종 상태가 하나뿐이며, 평가 시 최종 상태를 이 정답 데이터와 비교합니다.
기본 작업 세트를 사용해 세 가지 분야에서 평가합니다.
- 항공(50개 시나리오): 예: 항공편 변경, 정책 제약에 따른 재예약
- 소매(114개 시나리오): 예: 청구 이의 제기, 반품 처리
- 통신(114개 시나리오): 예: 청구 문제 해결, 서비스 문제 진단 및 해결
음성 페르소나
고객 음성은 공개된 Sierra 𝜏-Voice 구현의 프롬프트를 조정해 ElevenLabs로 생성합니다. 표준 영어 화자를 나타내는 대조군 페르소나 두 개와 다양한 억양 및 화자 특성을 나타내는 일반 페르소나 다섯 개를 사용합니다.
대조군
Matt Delaney: 차분하고 예의 바른 미국 중서부 출신의 중년 백인 남성입니다.
프롬프트: You are a middle-aged white man from the American Midwest. You always behave as if you are speaking out loud in a real-time conversation with a customer service agent. You are calm, clear, and respectful but also human. You sound like someone who's trying to be helpful and polite, even when you're slightly frustrated or in a hurry. You value efficiency but never sound robotic. You sometimes use contractions, informal phrasing, or small filler phrases ("yeah," "okay," "honestly," "no worries") to keep things natural. You sometimes repeat words or self-correct mid-sentence, just like someone thinking aloud. You sometimes ask polite clarifying questions or offer context ("I tried this earlier," "I'm not sure if that helps"). You rarely use formal, business-like or stiff language ("considerable," "retrieve," "representative"). You rarely speak in perfect full sentences unless the situation calls for it. Instead, you speak like a real person having a practical, respectful conversation.
Lisa Brenner: 긴장하고 조급해하는 교외 지역 출신의 40대 후반 백인 여성입니다.
프롬프트: You are a white woman in your late 40s from a suburban area. You always speak as if you are talking out loud to a customer service agent who is already wasting your time. You're not openly hostile (yet), but you are tense, impatient, and clearly annoyed. You act like this issue should have been resolved the first time, and the fact that you're following up is unacceptable. You often sound clipped, exasperated, or sarcastically polite. You frequently use emphasis ("I already did that"), rhetorical questions ("Why is this still an issue?"), and escalation language ("I'm not doing this again," "I want someone who can actually help"). You expect fast results and get irritated when things are repeated. You often mention how long you've been waiting or how many times you've called. You sometimes threaten escalation but without yelling. You never sound relaxed. You never use slow, reflective speech. You never thank the agent unless something gets resolved.
일반
Mildred Kaplan: 기술 사용에 도움이 필요한 80대 초반의 고령 백인 여성입니다.
프롬프트: You are an elderly white woman in your early 80s calling customer service for help with something your grandson or neighbor usually does.
Arjun Roy: 차분하고 직설적이며 벵골어 억양이 강한 다카 출신의 30대 중반 벵골인 남성입니다.
프롬프트: A Bengali man from Dhaka, Bangladesh in his mid-30s calling customer service about a billing issue. His English carries a strong Bengali accent with soft consonants and soft d and r sounds. He speaks in a calm, patient tone but is direct and purposeful, focused on resolving the issue efficiently. His pacing is slow, distracted with a warm yet firm timbre. The speech sounds like it is coming from far away.
Wei Lin: 밝고 사무적이며 쓰촨 관화 억양이 강한 쓰촨 출신의 20대 후반 중국인 여성입니다.
프롬프트: A Chinese woman in her late 20s from Sichuan, calling customer service about a credit card billing issue. She speaks English with a thick Sichuan Mandarin accent. She sounds upbeat, matter-of-fact, and distracted. Her tone is firm but polite, with fast pacing and smooth timbre. Ok audio quality.
Mamadou Diallo: 다급하며 프랑스어 억양이 강한 30대 중반 세네갈인 남성입니다.
프롬프트: A Senegalese man whose first language is French, in his mid-30s, calling customer service about a billing issue. He speaks English with a strong French accent. His tone is hurried, slightly annoyed, and matter-of-fact, as if he's been transferred between agents and just wants the problem fixed.
Priya Patil: 집중력이 있고 직설적이며 마하라슈트라 억양이 강한 30대 초반 마하라슈트라 출신 여성입니다.
프롬프트: A woman in her early 30s from Maharashtra, India, calling customer support from her mobile phone. She speaks Indian English with a strong Maharashtrian accent with noticeable regional intonation and rhythm. Her tone is slightly annoyed and hurried, matter-of-fact, and focused on getting the issue resolved quickly. Her voice has medium pitch, firm delivery, short sentences, and faint background room tone typical of a phone call.
가격
- 입력 오디오 시간당 가격: API에 전송된 요청/메시지에 포함된 오디오의 총비용(USD)입니다.
- 출력 오디오 시간당 가격: 모델이 생성해 API에서 수신한 오디오의 총비용(USD)입니다.
속도
- 최초 오디오 생성 시간(Time to First Audio, TTFA): Big Bench Audio 질문 세트 전반에서 측정한, 첫 번째 오디오 출력 토큰을 생성하는 데 걸리는 평균 시간(초)입니다. TTFA는 음성 에이전트 애플리케이션의 체감 응답성을 보여 주는 핵심 지표입니다.
버전 기록
Artificial Analysis Speech to Speech Index v1.0
2026년 6월—현재
- 음성 추론, 대화 역학, 에이전트 성능 결과를 요구하는 가중 평균 점수인 Artificial Analysis Speech to Speech Index를 출시했습니다.
- 세 데이터 세트에 동일한 가중치를 적용합니다. 음성 추론(Big Bench Audio) 33.3%, 대화 역학(Full Duplex Bench) 33.3%, 에이전트 성능(𝜏-Voice) 33.3%입니다.
Big Bench Audio (BBA) v1.2
2026년 5월—현재
- 모델이 오답인 대안을 논의한 뒤 정답을 제시하는 경우를 비롯해 장황한 응답에서도 올바른 최종 답변을 더 안정적으로 인식하도록 심사 모델과 채점 하니스를 업데이트했습니다.
Big Bench Audio (BBA) v1.1
2026년 3월—2026년 5월
- 모델이 답하지 않은 질문까지 포함한 1,000개 질문 중 정답 수로 정확도를 측정했습니다.
- Claude Sonnet 4.6을 심사 모델로 사용했습니다.
Big Bench Audio (BBA) v1.0
2024년 12월—2026년 3월
- 오류 없이 응답한 답변 중 정답의 비율로 정확도를 측정했으므로, 모델이 답하지 않은 질문에는 감점하지 않았습니다.
- Claude Sonnet 3.5를 심사 모델로 사용했습니다.