다국어 AI 모델 벤치마크 언어별 주요 LLM 비교

Global-MMLU-Lite 벤치마크를 포함한 Artificial Analysis Multilingual Index에서 주요 대규모 언어 모델(LLM)이 여러 언어에 걸쳐 어떤 성능을 보이는지 살펴보세요. 언어와 모델로 필터링하고 정확도, 속도, 비용의 상충 관계를 확인해 다국어 사용 사례에 가장 적합한 LLM을 찾을 수 있습니다.

데이터 세트와 방법론에 관한 자세한 내용은 FAQ 페이지에서 확인하세요.

개요

Artificial Analysis Multilingual Index

Higher is better

언어별 Multilingual Index(정규화)

테스트한 모든 모델에서 언어별 점수를 정규화합니다. 초록색은 해당 언어의 최고 점수, 빨간색은 최저 점수를 나타냅니다.

Multilingual Index

Multilingual Index: 전체 언어 평균

Artificial Analysis Multilingual Index · Average across all languages · Higher is better

Multilingual Index: 평균과 출력 속도

Artificial Analysis Multilingual Index · Output speed: output tokens per second
Most attractive quadrant

Multilingual Index: 평균과 가격

Artificial Analysis Multilingual Index · Average across all languages
Most attractive quadrant

Global-MMLU-Lite

다국어 Global-MMLU-Lite: 평균

Average across all languages · Higher is better

가격

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens)

속도 및 지연 시간

Output Speed

Output tokens per second · Higher is better

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time

End-to-End Response Time

Seconds to output 500 tokens, including reasoning model 'thinking' time · Lower is better