DeepSeek 模型:智能、性能与价格
Artificial Analysis Intelligence Index
Artificial Analysis 已对 DeepSeek 的 36 个模型进行基准测试。下方对这些模型的关键指标进行了比较。
- 智能方面,DeepSeek 得分最高的模型是 DeepSeek V4.1 Flash (Reasoning, Max Effort)(39)。
- 输出速度方面,最快的模型是 DeepSeek V4.1 Flash (Non-Reasoning)(239 token/秒)。
- 延迟方面,DeepSeek V4.1 Flash (Non-Reasoning)(1.12 秒) 的首个答案 token 等待时间最短。
- 价格方面,DeepSeek V4.1 Flash (Non-Reasoning)($0.15) 的单任务成本最低。不同模型的价格差异最高可达 4.6 倍。
智能
Artificial Analysis Intelligence Index
Intelligence Index vs. Cost per Intelligence Index Task
成本
Cost per Intelligence Index Task
速度与延迟
Output Speed
能力得分
能力指数
Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better
Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better
Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AA-LCR v1.1, GDP.pdf, AutomationBench-AA · Higher is better
Incorporates 6 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, MLCR-AA, Humanity's Last Exam, AutomationBench-AA · Higher is better
Incorporates 6 evaluations: AA-Omniscience, Humanity's Last Exam, CritPt, GDPval-AA v2.1, AA-Briefcase v1.1, Terminal-Bench 4.0 · Higher is better
Incorporates 5 evaluations: AA-Omniscience, Humanity's Last Exam, GDPval-AA v2.1, AA-Briefcase v1.1, AA-LCR v1.1 · Higher is better
DeepSeek 的所有发布
DeepSeek V4.1 Flash
2 个模型
39
最高智能得分
DeepSeek V4 Flash Vision
1 个模型
35
最高智能得分
DeepSeek V4 Pro 0813
1 个模型
36
最高智能得分
DeepSeek V4 Flash 0731
1 个模型
34
最高智能得分
DeepSeek V4 Flash 0420
3 个模型
26
最高智能得分
DeepSeek V4 Pro 0424
3 个模型
30
最高智能得分
DeepSeek V3.2
2 个模型
21
最高智能得分
DeepSeek V3.2 Speciale
1 个模型
14
最高智能得分
DeepSeek V3.2 Exp
2 个模型
17
最高智能得分
DeepSeek V3.1 Terminus
2 个模型
15
最高智能得分
DeepSeek V3.1
2 个模型
14
最高智能得分
DeepSeek R1 0528 Qwen3 8B
1 个模型
8
最高智能得分
DeepSeek R1 0528 (May '25)
1 个模型
13
最高智能得分
DeepSeek V3 0324
1 个模型
10
最高智能得分
DeepSeek R1 (Jan '25)
1 个模型
11
最高智能得分
DeepSeek R1 Distill Llama 70B
1 个模型
8
最高智能得分
DeepSeek R1 Distill Llama 8B
1 个模型
6
最高智能得分
DeepSeek R1 Distill Qwen 1.5B
1 个模型
6
最高智能得分
DeepSeek R1 Distill Qwen 14B
1 个模型
8
最高智能得分
DeepSeek R1 Distill Qwen 32B
1 个模型
8
最高智能得分
DeepSeek V3 (Dec '24)
1 个模型
8
最高智能得分
DeepSeek-V2.5 (Dec '24)
1 个模型
7
最高智能得分
DeepSeek-V2.5
1 个模型
7
最高智能得分
DeepSeek-Coder-V2
1 个模型
6
最高智能得分
DeepSeek Coder V2 Lite Instruct
1 个模型
5
最高智能得分
DeepSeek-V2-Chat
1 个模型
6
最高智能得分
DeepSeek LLM 67B Chat (V1)
1 个模型
5
最高智能得分
更多详情
权重 | 服务商基准测试 | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash (Reasoning, Max Effort) | 39 | 552B 推理时启用 16B 个参数 | 1M | US$0.2 | 221 | +15 | |||
| DeepSeek V4 Pro 0813 (Reasoning, Max Effort) | 36 | 1.6T 推理时启用 49B 个参数 | 1M | US$0.7 | 104 | +8 | |||
| DeepSeek V4 Flash Vision (Reasoning, Max Effort) | 35 | 284B 推理时启用 13B 个参数 | 1M | US$0.2 | 223 | 不可用 | +2 | ||
| DeepSeek V4 Flash 0731 (Reasoning, Max Effort) | 34 | 284B 推理时启用 13B 个参数 | 1M | US$0.2 | 227 | +14 | |||
| DeepSeek V4 Pro 0424 (Reasoning, Max Effort) | 30 | 1.6T 推理时启用 49B 个参数 | 1M | US$0.2 | 111 | +7 | |||
| DeepSeek V4 Pro 0424 (Reasoning, High Effort) | 30 | 1.6T 推理时启用 49B 个参数 | 1M | US$0.2 | 106 | +5 | |||
| DeepSeek V4 Flash 0420 (Reasoning, High Effort) | 26 | 284B 推理时启用 13B 个参数 | 1M | US$0.1 | - | +4 | |||
| DeepSeek V4.1 Flash (Non-Reasoning) | 25 | 552B 推理时启用 16B 个参数 | 1M | US$0.2 | 239 | ||||
| DeepSeek V4 Flash 0420 (Reasoning, Max Effort) | 24 | 284B 推理时启用 13B 个参数 | 1M | US$0.1 | - | +2 | |||
| DeepSeek V3.2 (Reasoning) | 21 | 685B 推理时启用 37B 个参数 | 128k | US$0.1 | - | +5 | |||
| DeepSeek V4 Pro 0424 (Non-reasoning) | 21 | 1.6T 推理时启用 49B 个参数 | 1M | US$0.2 | 99 | +3 | |||
| DeepSeek V4 Flash 0420 (Non-reasoning) | 19 | 284B 推理时启用 13B 个参数 | 1M | US$0.1 | - | ||||
| DeepSeek V3.2 Exp (Reasoning) | 17 | 685B 推理时启用 37B 个参数 | 128k | US$0.1 | - | ||||
| DeepSeek V3.2 (Non-reasoning) | 16 | 685B 推理时启用 37B 个参数 | 128k | US$0.3 | - | +8 | |||
| DeepSeek V3.1 Terminus (Reasoning) | 15 | 685B 推理时启用 37B 个参数 | 128k | US$1.7 | - | ||||
| DeepSeek V3.2 Speciale | 14 | 685B 推理时启用 37B 个参数 | 128k | - | - | - | |||
| DeepSeek V3.1 Terminus (Non-reasoning) | 14 | 685B 推理时启用 37B 个参数 | 128k | US$0.2 | - | ||||
| DeepSeek V3.2 Exp (Non-reasoning) | 14 | 685B 推理时启用 37B 个参数 | 128k | US$0.1 | - | ||||
| DeepSeek V3.1 (Non-reasoning) | 14 | 685B 推理时启用 37B 个参数 | 128k | US$0.7 | - | +5 | |||
| DeepSeek V3.1 (Reasoning) | 13 | 685B 推理时启用 37B 个参数 | 128k | US$0.7 | - | ||||
| DeepSeek R1 0528 (May '25) | 13 | 685B 推理时启用 37B 个参数 | 128k | US$1.5 | - | +2 | |||
| DeepSeek R1 (Jan '25) | 11 | 685B 推理时启用 37B 个参数 | 128k | US$2.2 | - | ||||
| DeepSeek V3 0324 | 10 | 671B 推理时启用 37B 个参数 | 128k | US$0.8 | - | ||||
| DeepSeek V3 (Dec '24) | 8 | 671B 推理时启用 37B 个参数 | 128k | US$0.4 | - | ||||
| DeepSeek R1 Distill Qwen 32B | 8 | 32B | 128k | - | - | - | |||
| DeepSeek R1 0528 Qwen3 8B | 8 | 8.2B | 33k | - | - | - | |||
| DeepSeek R1 Distill Llama 70B | 8 | 70B | 128k | US$0.7 | - | ||||
| DeepSeek R1 Distill Qwen 14B | 8 | 14B | 128k | - | - | - | |||
| DeepSeek-V2.5 (Dec '24) | 7 | 236B 推理时启用 21B 个参数 | 128k | - | - | - | |||
| DeepSeek-V2.5 | 7 | 236B 推理时启用 21B 个参数 | 128k | - | - | - | |||
| DeepSeek R1 Distill Llama 8B | 6 | 8B | 128k | - | - | - | |||
| DeepSeek-Coder-V2 | 6 | 236B 推理时启用 21B 个参数 | 128k | - | - | - | |||
| DeepSeek R1 Distill Qwen 1.5B | 6 | 1.5B | 128k | - | - | - | |||
| DeepSeek-V2-Chat | 6 | 236B 推理时启用 21B 个参数 | 128k | - | - | - | |||
| DeepSeek Coder V2 Lite Instruct | 5 | 16B 推理时启用 2.4B 个参数 | 128k | - | - | - | |||
| DeepSeek LLM 67B Chat (V1) | 5 | 7B | 4k | - | - | - |