多语言 AI 模型基准测试 按语言比较领先 LLM

探索领先大语言模型(LLM)在 Artificial Analysis 多语言指数中的多语言表现,其中包括 Global-MMLU-Lite 基准测试。按语言和模型筛选,查看准确率、速度与成本之间的权衡,为您的多语言使用场景寻找最佳 LLM。

如需了解数据集和方法论详情,请参阅常见问题页面

概述

Artificial Analysis 多语言指数

Higher is better

各语言的多语言指数(归一化)

分数按语言在所有受测模型中归一化,其中绿色表示该语言的最高分,红色表示最低分。

多语言指数

多语言指数:所有语言的平均分

Artificial Analysis Multilingual Index · Average across all languages · Higher is better

多语言指数:平均分与输出速度

Artificial Analysis Multilingual Index · Output speed: output tokens per second
Most attractive quadrant

多语言指数:平均分与价格

Artificial Analysis Multilingual Index · Average across all languages
Most attractive quadrant

Global-MMLU-Lite

多语言 Global-MMLU-Lite:平均分

Average across all languages · Higher is better

价格

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens)

速度与延迟

Output Speed

Output tokens per second · Higher is better

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time

End-to-End Response Time

Seconds to output 500 tokens, including reasoning model 'thinking' time · Lower is better