多言語AIモデルベンチマーク 言語別に主要LLMを比較

Global-MMLU-Liteベンチマークを含むArtificial Analysis Multilingual Indexで、主要な大規模言語モデル(LLM)が複数の言語にわたって示す性能を確認できます。言語とモデルで絞り込み、精度、速度、費用のトレードオフを確認し、多言語のユースケースに最適なLLMを見つけられます。

データセットと方法論の詳細は、よくある質問のページをご覧ください。

概要

Artificial Analysis Multilingual Index

Higher is better

言語別Multilingual Index(正規化)

テストした全モデルを対象に言語ごとにスコアを正規化しています。緑はその言語で最高のスコア、赤は最低のスコアを示します。

Multilingual Index

Multilingual Index:全言語の平均

Artificial Analysis Multilingual Index · Average across all languages · Higher is better

Multilingual Index:平均と出力速度

Artificial Analysis Multilingual Index · Output speed: output tokens per second
Most attractive quadrant

Multilingual Index:平均と料金

Artificial Analysis Multilingual Index · Average across all languages
Most attractive quadrant

Global-MMLU-Lite

多言語Global-MMLU-Lite:平均

Average across all languages · Higher is better

料金

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens)

速度と遅延

Output Speed

Output tokens per second · Higher is better

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time

End-to-End Response Time

Seconds to output 500 tokens, including reasoning model 'thinking' time · Lower is better