英语——AI 模型基准测试 比较多语言 LLM 表现

适用于英语任务的前 5 个 AI 模型是 Claude Opus 4.6 (max)、Gemini 3.1 Pro Preview、Gemini 3 Flash、Claude Opus 4.5 (Non-reasoning) 和 GPT-5.1 (high)。它们在 Artificial Analysis 多语言指数中取得了最高的英语推理得分。

如需比较所有支持语言的表现,请查看完整的多语言 AI 模型基准测试页面。

🇬🇧 适用于英语的顶尖模型

#1
Claude Opus 4.6 (Adaptive Reasoning, Max Effort)Claude Opus 4.6 (max)
95#2
Gemini 3.1 Pro PreviewGemini 3.1 Pro Preview
95#3
Gemini 3 Flash Preview (Reasoning)Gemini 3 Flash
95#4
Claude Opus 4.5 (Non-reasoning)Claude Opus 4.5 (Non-reasoning)
94#5
GPT-5.1 (high)GPT-5.1 (high)
94

亮点

Multilingual Index: English · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens (blended) · Lower is better

多语言指数

多语言指数:英语

Artificial Analysis Multilingual Index · Higher is better

多语言指数:英语与价格

Artificial Analysis Multilingual Index · USD per 1M tokens (blended)
Most attractive quadrant

多语言指数:英语与输出速度

Artificial Analysis Multilingual Index · Output speed: output tokens per second
Most attractive quadrant

多语言指数:英语与上下文窗口

Artificial Analysis Multilingual Index · Context window: tokens limit
Most attractive quadrant

Global-MMLU-Lite

多语言 Global-MMLU-Lite:英语

Multilingual Global-MMLU-Lite · Higher is better

价格

Pricing: Cache Hit, Input, and Output

Price (USD per M Tokens)

速度与延迟

Output Speed

Output tokens per second · Higher is better

Latency: Time To First Answer Token

Seconds to first answer token received · Accounts for reasoning model 'thinking' time

End-to-End Response Time

Seconds to output 500 tokens, including reasoning model 'thinking' time · Lower is better