DeepInfra:モデルの知能、性能、料金

この分析は、ユースケースに最適なDeepInfra提供モデルを選ぶための参考情報です。
最高の知能
UpdatedIntelligence Index
モデル合計:108件
最速
出力速度
モデル合計:108件
最安料金
ブレンド料金(100万トークンあたり)
モデル合計:108件
DeepInfraは108モデルを提供しており、知能、性能、料金の特性はそれぞれ異なります。 以下でモデル間の主要指標を比較します。
- DeepInfraで知能が上位のモデルはGLM-5.3 (max)(45)、GLM-5.3-Flash(42)、Qwen3.8 2.4T A95B(40)です。
- 出力速度が最も速いモデルはInkling Small(244 t/s)、gpt-oss-120b (high) (Turbo)(230 t/s)、Granite 4.2 3B(214 t/s)です。
- 遅延では、Llama 4 Maverick (FP8)(0.50秒)、Qwen3 30B (Non-reasoning) (FP8)(0.51秒)、Qwen3 Coder 480B (Turbo, FP4)(0.64秒)の最初の回答トークンまでの時間が最短です。
- 料金では、Llama 3.1 8B (Turbo, FP8)($0.02)、Llama 3.1 8B($0.02)、Granite 4.2 3B($0.02)の100万トークンあたりのブレンド料金が最安です。
- DeepInfraで最大のコンテキストウィンドウに対応するモデルはGLM-5.2 (max) (FP4)(1M)、DeepSeek V4 Pro (max) (FP4)(1M)、DeepSeek V4 Flash (high) (FP4)(1M)です。
ハイライト
知能評価
Artificial Analysis Intelligence Index
Intelligence Evaluations
Agentic knowledge work, (Elo-500)/2000
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Quantitative analysis on spreadsheets & documents
Instruction following
Agentic tool use
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
Intelligence Index vs. Price
コンテキストウィンドウ
Context Window
料金
Intelligence Index vs. Price
性能の概要
Output Speed vs. Price
速度
出力速度(1秒あたりのトークン数)で測定
Output Speed
遅延
最初のトークンまでの時間(秒)で測定
Latency: Time To First Answer Token
エンドツーエンド応答時間
Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed
End-to-End Response Time vs. Price
詳細分析 | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
GLM-5.3 (max) | 1.05M | オープン | 45 | $2.16 | 141 | 1.04 | 18.82 | 14.22 | |||
GLM-5.3-Flash | 1M | オープン | 42 | $0.43 | 31 | 1.36 | 81.71 | 64.28 | |||
Qwen3.8 2.4T A95B | 262k | オープン | 40 | $2.89 | 89 | 1.48 | 29.69 | 22.56 | |||
DeepSeek V4 Pro 0813 (max) | 1M | オープン | 36 | $1.00 | 114 | 0.80 | 22.66 | 17.49 | |||
DeepSeek V4 Flash 0731 (max) | 1M | オープン | 35 | $0.10 | 57 | 0.94 | 45.09 | 35.32 | |||
GLM-5.2 (max) (FP4) | 1.05M | オープン | 34 | $0.65 | 63 | 1.05 | 40.91 | 31.89 | |||
Qwen3.8 27B (xhigh) | 262k | オープン | 34 | $0.65 | 36 | 1.02 | 70.84 | 55.86 | |||
DeepSeek V4 Pro (max) (FP4) | 1.05M | オープン | 31 | $1.19 | 52 | 1.09 | 94.64 | 83.95 | |||
DeepSeek V4 Pro (high) (FP4) | 65.5k | オープン | 30* | -- | 56 | 1.50 | 45.91 | 35.50 | |||
MiniMax-M3 | 524k | オープン | 30 | $0.44 | 24 | 0.96 | 106.39 | 84.35 | |||
GLM-5 (FP4) | 203k | オープン | 28* | -- | 70 | 0.98 | 52.34 | 44.24 | |||
Kimi K2.6 (FP4) | 262k | オープン | 27 | $0.51 | 45 | 1.43 | 110.34 | 97.91 | |||
GLM-5.1 (FP4) | 203k | オープン | 26 | $0.69 | 31 | 0.83 | 137.60 | 120.83 | |||
MiMo-V2.5-Pro | 65.5k | オープン | 26 | $0.43 | 37 | 1.11 | 68.77 | 54.13 | |||
Kimi K2.7 Code | 262k | オープン | 26 | $0.39 | 41 | 0.95 | 67.92 | 54.70 | |||
Inkling Small | 524k | オープン | 26 | $0.12 | 244 | 0.98 | 11.24 | 8.21 | |||
Hy3 (FP8) | 262k | オープン | 26 | $0.07 | 60 | 1.28 | 42.76 | 33.19 | |||
Inkling (FP8) | 131k | オープン | 26 | $0.81 | 204 | 0.56 | 12.82 | 9.81 | |||
DeepSeek V4 Flash (high) (FP4) | 1.05M | オープン | 25 | $0.10 | 53 | 0.89 | 34.02 | 23.61 | |||
DeepSeek V4 Flash (max) (FP4) | 1.05M | オープン | 25 | $0.09 | 42 | 0.95 | 145.74 | 132.95 | |||
GLM-5.1 (Non-reasoning) (FP4) | 203k | オープン | 24* | -- | 34 | 1.05 | 15.84 | -- | |||
Kimi K2.6 (Non-reasoning) (FP4) | 262k | オープン | 24* | -- | 36 | 1.49 | 15.39 | -- | |||
Kimi K2.5 | 262k | オープン | 23* | -- | 47 | 1.34 | 75.25 | 63.25 | |||
Nemotron 3 Ultra | 262k | オープン | 23 | $0.47 | 143 | 3.48 | 22.89 | 15.92 | |||
Nemotron 3 Ultra BF16 | 262k | オープン | 23 | $0.92 | 144 | 3.26 | 22.50 | 15.77 | |||
Qwen3.5 27B (FP8) | 262k | オープン | 23* | -- | 70 | 0.90 | 36.81 | 28.73 | |||
MiniMax-M2.5 (FP8) | 197k | オープン | 23* | -- | 22 | 3.30 | 116.26 | 90.36 | |||
MiMo-V2.5 | 262k | オープン | 22 | $0.02 | 29 | 1.57 | 87.29 | 68.58 | |||
GLM-4.7 (FP4) | 203k | オープン | 22* | -- | 18 | 1.52 | 141.69 | 112.14 | |||
Qwen3.6 27B FP8 | 262k | オープン | 22 | $0.37 | 67 | 0.94 | 93.57 | 85.14 | |||
GLM-5 (Non-reasoning) (FP8) | 203k | オープン | 22* | -- | 54 | 1.33 | 10.64 | -- | |||
DeepSeek V3.2 (FP4) | 164k | オープン | 21* | -- | 27 | 1.67 | 92.95 | 73.03 | |||
Qwen3.5 397B A17B (Non-reasoning) (FP8) | 262k | オープン | 21* | -- | 41 | 0.97 | 13.07 | -- | |||
Ling 3.0 Flash | 131k | オープン | 21 | $0.03 | 61 | 0.63 | 41.39 | 32.61 | |||
Qwen3.6 27B (Non-reasoning) FP8 | 262k | オープン | 20* | -- | 61 | 0.91 | 9.07 | -- | |||
Step 3.7 Flash | 256k | オープン | 19* | -- | 168 | 0.66 | 15.52 | 11.89 | |||
Qwen3.5 27B (Non-reasoning) FP8 | 262k | オープン | 19* | -- | 67 | 0.92 | 8.41 | -- | |||
Qwen3.5 35B A3B (FP8) | 262k | オープン | 19* | -- | 159 | 0.80 | 16.48 | 12.54 | |||
Qwen3.5 397B A17B (FP8) | 262k | オープン | 19 | $0.23 | 39 | 0.97 | 94.63 | 80.95 | |||
Qwen3.6 35B A3B (FP8) | 262k | オープン | 19 | $0.14 | 97 | 0.70 | 61.76 | 55.88 | |||
GLM-4.6 (FP4) | 203k | オープン | 19* | -- | 41 | 0.75 | 62.48 | 49.38 | |||
MiMo-V2.5-Pro (Non-reasoning) | 65.5k | オープン | 18* | -- | 32 | 1.16 | 16.83 | -- | |||
Muse Glimmer (high) | 131k | オープン | 18 | $0.05 | 145 | 0.94 | 18.14 | 13.76 | |||
Qwen3.5 122B A10B (Non-reasoning) (FP4) | 262k | オープン | 18* | -- | 143 | 0.92 | 4.42 | -- | |||
GLM-4.7 (Non-reasoning) (FP4) | 203k | オープン | 17* | -- | 18 | 1.10 | 28.40 | -- | |||
Gemma 4 26B A4B (FP8) | 262k | オープン | 17* | -- | 30 | 0.87 | 84.86 | 67.19 | |||
Qwen3.5 122B A10B (FP4) | 262k | オープン | 16 | $0.14 | 152 | 0.79 | 17.27 | 13.18 | |||
DeepSeek V3.2 (Non-reasoning) | 164k | オープン | 16* | -- | 32 | 1.35 | 16.89 | -- | |||
Gemma 4 31B | 262k | オープン | 15 | $0.02 | 15 | 2.14 | 149.38 | 114.32 | |||
Qwen3.6 35B A3B (Non-reasoning) (FP8) | 262k | オープン | 15* | -- | 108 | 0.69 | 5.33 | -- | |||
Qwen3.5 35B A3B (Non-reasoning) FP8 | 262k | オープン | 15* | -- | 154 | 0.72 | 3.97 | -- | |||
GLM-4.7-Flash | 203k | オープン | 15* | -- | 35 | 1.56 | 72.92 | 57.09 | |||
Granite 4.2 30B | 131k | オープン | 15* | -- | 77 | 0.81 | 33.19 | 25.90 | |||
DeepSeek V3.1 Terminus (Non-reasoning) (FP4) | 164k | オープン | 14* | -- | 60 | 0.97 | 9.35 | -- | |||
Gemma 4 31B (Non-reasoning) (FP8) | 262k | オープン | 14* | -- | 25 | 1.45 | 21.60 | -- | |||
DeepSeek V3.1 (Non-reasoning) (FP4) | 164k | オープン | 14* | -- | 12 | 1.00 | 42.55 | -- | |||
Nemotron 3.5 Lightning (NVFP4) | 262k | オープン | 14 | $0.12 | -- | -- | -- | -- | |||
Nemotron 3 Super | 262k | オープン | 14 | $0.48 | 87 | 14.99 | 43.62 | 22.91 | |||
Gemma 4 26B A4B (Non-reasoning) (FP8) | 262k | オープン | 13* | -- | 23 | 0.87 | 22.47 | -- | |||
Qwen3.5 4B (FP8) | 262k | オープン | 13* | -- | 29 | 0.70 | 85.83 | 68.10 | |||
DeepSeek R1 0528 | 164k | オープン | 13* | -- | 24 | 0.97 | 105.07 | 83.28 | |||
Qwen3 235B A22B 2507 (FP8) | 262k | オープン | 13 | -- | 114 | 0.72 | 22.72 | 17.60 | |||
gpt-oss-120b (high) (Turbo) | 131k | オープン | 12 | $0.11 | 230 | 0.72 | 11.58 | 8.69 | |||
gpt-oss-120b (high) | 131k | オープン | 12 | $0.03 | 53 | 0.61 | 47.75 | 37.71 | |||
Qwen3 235B 2507 (Non-reasoning) (FP8) | 262k | オープン | 12* | -- | 16 | 0.67 | 32.20 | -- | |||
Qwen3 Coder 480B (Turbo, FP4) | 262k | オープン | 12* | -- | 79 | 0.64 | 6.93 | -- | |||
Granite 4.2 8B | 131k | オープン | 12 | $0.02 | 76 | 0.62 | 33.54 | 26.34 | |||
Qwen3.5 4B (Non-reasoning) FP8 | 262k | オープン | 11* | -- | 23 | 0.69 | 21.98 | -- | |||
gpt-oss-120b (low) | 131k | オープン | 10* | -- | 55 | 0.63 | 46.38 | 36.60 | |||
DeepSeek V3 0324 (FP4) | 164k | オープン | 10 | $0.02 | 46 | 1.16 | 12.14 | -- | |||
Qwen3 Next 80B A3B | 262k | オープン | 10* | -- | 160 | 0.68 | 3.80 | -- | |||
Llama 4 Maverick (FP8) | 1.05M | オープン | 9 | $0.03 | 74 | 0.50 | 7.25 | -- | |||
Granite 4.2 3B | 131k | オープン | 9 | $0.01 | 214 | 0.42 | 12.11 | 9.35 | |||
gpt-oss-20b (high) | 131k | オープン | 9 | $0.01 | 117 | 0.44 | 21.85 | 17.13 | |||
Llama Nemotron Super 49B v1.5 | 131k | オープン | 9* | -- | 95 | 4.25 | 30.66 | 21.13 | |||
Gemma 4 E4B | 262k | オープン | 9* | -- | 44 | 0.82 | 57.83 | 45.61 | |||
Nemotron 3 Nano | 262k | オープン | 9 | $0.02 | 97 | 5.80 | 31.45 | 20.52 | |||
DeepSeek V3 (Dec) | 164k | オープン | 8 | $0.02 | 29 | 0.70 | 17.85 | -- | |||
DeepSeek R1 Distill Llama 70B | 131k | オープン | 8* | -- | 25 | 1.00 | 101.29 | 80.23 | |||
Qwen2.5 72B | 32.8k | オープン | 8* | -- | 29 | 2.33 | 19.47 | -- | |||
Llama 3.3 70B (Turbo, FP8) | 131k | オープン | 8* | -- | 16 | 2.05 | 34.11 | -- | |||
Qwen3 30B (FP8) | 41k | オープン | 8* | -- | 157 | 0.51 | 16.46 | 12.76 | |||
NVIDIA Nemotron Nano 12B v2 VL (FP8) | 131k | オープン | 7* | -- | 137 | 3.82 | 22.01 | 14.55 | |||
Gemma 4 E4B (Non-reasoning) | 262k | オープン | 7* | -- | 45 | 0.78 | 11.85 | -- | |||
NVIDIA Nemotron Nano 9B V2 | 131k | オープン | 7* | -- | 95 | 7.88 | 34.25 | 21.10 | |||
Mistral Small 3.1 | 128k | オープン | 7 | $0.02 | 36 | 0.94 | 14.67 | -- | |||
Llama Nemotron Super 49B v1.5 (Non-reasoning) | 131k | オープン | 7* | -- | 144 | 3.02 | 6.49 | -- | |||
Qwen3 32B (Non-reasoning) (FP8) | 41k | オープン | 7* | -- | 27 | 1.43 | 19.72 | -- | |||
Qwen3 32B (FP8) | 41k | オープン | 7 | -- | 29 | 1.47 | 88.80 | 69.86 | |||
Mistral Small 3.2 (FP8) | 128k | オープン | 7 | $0.11 | 37 | 0.92 | 14.36 | -- | |||
Llama 3.1 Nemotron 70B | 131k | オープン | 7* | -- | 126 | 3.22 | 7.18 | -- | |||
Llama 3.1 8B (Turbo, FP8) | 131k | オープン | 7* | -- | 23 | 0.95 | 22.25 | -- | |||
Llama 3.1 8B | 131k | オープン | 7* | -- | 20 | 0.97 | 25.69 | -- | |||
Nemotron 3 Nano (Non-reasoning) | 262k | オープン | 7* | -- | 103 | 0.64 | 5.51 | -- | |||
NVIDIA Nemotron Nano 9B V2 (Non-reasoning) | 131k | オープン | 7* | -- | 102 | 7.39 | 12.29 | -- | |||
Qwen3 14B (Non-reasoning) (FP8) | 41k | オープン | 7* | -- | 64 | 0.80 | 8.59 | -- | |||
Mistral Small 3 | 32.8k | オープン | 7* | -- | 53 | 0.83 | 10.27 | -- | |||
Qwen3 30B (Non-reasoning) (FP8) | 41k | オープン | 7* | -- | 164 | 0.51 | 3.56 | -- | |||
Llama 3.1 70B | 131k | オープン | 7* | -- | 35 | 1.90 | 16.11 | -- | |||
Llama 3.1 70B (Turbo, FP8) | 131k | オープン | 7* | -- | 36 | 1.94 | 15.80 | -- | |||
Llama 4 Scout | 328k | オープン | 6 | $0.06 | 22 | 0.91 | 23.55 | -- | |||
Qwen3 14B (FP8) | 32.8k | オープン | 6 | -- | 57 | 0.99 | 45.03 | 35.23 | |||
Hermes 3 - Llama-3.1 70B | 131k | オープン | 6* | -- | 31 | 2.13 | 18.17 | -- | |||
Phi-4 | 16.4k | オープン | 6* | -- | 56 | 0.93 | 9.82 | -- | |||
NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) (FP8) | 131k | オープン | 6* | -- | 128 | 4.58 | 8.48 | -- | |||
Llama 3.2 11B (Vision) | 131k | オープン | 5* | -- | 22 | 1.27 | 23.75 | -- | |||
Gemma 3 27B | 131k | オープン | 5 | $0.16 | 25 | 1.37 | 21.27 | -- | |||
Gemma 3 4B | 131k | オープン | 5* | -- | 22 | 1.52 | 23.96 | -- | |||
Llama 3 8B | 8.19k | オープン | 5* | -- | -- | -- | -- | -- | |||
Gemma 3 12B | 131k | オープン | 4 | $0.13 | 39 | 1.04 | 13.74 | -- | |||
主要な定義
よくある質問
DeepInfraに関するよくある質問
DeepInfraが提供し、当社が追跡しているモデルは108モデルです:GLM-5.3 (max)、GLM-5.3-Flash、Qwen3.8 2.4T A95B、DeepSeek V4 Pro 0813 (max)、DeepSeek V4 Flash 0731 (max)、GLM-5.2 (max) (FP4)、Qwen3.8 27B (xhigh)、Kimi K2.6 (FP4)、DeepSeek V4 Pro (max) (FP4)、DeepSeek V4 Pro (high) (FP4)、MiniMax-M3、GLM-5 (FP4)、GLM-5.1 (FP4)、MiMo-V2.5-Pro、Kimi K2.7 Code、Inkling Small、Hy3 (FP8)、Inkling (FP8)、DeepSeek V4 Flash (high) (FP4)、DeepSeek V4 Flash (max) (FP4)、GLM-5.1 (Non-reasoning) (FP4)、Kimi K2.6 (Non-reasoning) (FP4)、Kimi K2.5、Nemotron 3 Ultra、Nemotron 3 Ultra BF16、Qwen3.5 27B (FP8)、MiniMax-M2.5 (FP8)、MiMo-V2.5、GLM-4.7 (FP4)、Qwen3.6 27B FP8、GLM-5 (Non-reasoning) (FP8)、DeepSeek V3.2 (FP4)、Qwen3.5 397B A17B (Non-reasoning) (FP8)、Ling 3.0 Flash、Qwen3.6 27B (Non-reasoning) FP8、Step 3.7 Flash、Qwen3.5 27B (Non-reasoning) FP8、Qwen3.5 35B A3B (FP8)、Qwen3.5 397B A17B (FP8)、Qwen3.6 35B A3B (FP8)、GLM-4.6 (FP4)、MiMo-V2.5-Pro (Non-reasoning)、Muse Glimmer (high)、Qwen3.5 122B A10B (Non-reasoning) (FP4)、GLM-4.7 (Non-reasoning) (FP4)、Gemma 4 26B A4B (FP8)、Qwen3.5 122B A10B (FP4)、DeepSeek V3.2 (Non-reasoning)、Gemma 4 31B、Qwen3.6 35B A3B (Non-reasoning) (FP8)、Qwen3.5 35B A3B (Non-reasoning) FP8、GLM-4.7-Flash、Granite 4.2 30B、DeepSeek V3.1 Terminus (Non-reasoning) (FP4)、Gemma 4 31B (Non-reasoning) (FP8)、DeepSeek V3.1 (Non-reasoning) (FP4)、Nemotron 3 Super、Gemma 4 26B A4B (Non-reasoning) (FP8)、Qwen3.5 4B (FP8)、DeepSeek R1 0528、Qwen3 235B A22B 2507 (FP8)、gpt-oss-120b (high) (Turbo)、gpt-oss-120b (high)、Qwen3 235B 2507 (Non-reasoning) (FP8)、Qwen3 Coder 480B (Turbo, FP4)、Granite 4.2 8B、Qwen3.5 4B (Non-reasoning) FP8、gpt-oss-120b (low)、DeepSeek V3 0324 (FP4)、Qwen3 Next 80B A3B、Llama 4 Maverick (FP8)、Granite 4.2 3B、gpt-oss-20b (high)、Llama Nemotron Super 49B v1.5、Gemma 4 E4B、Nemotron 3 Nano、DeepSeek V3 (Dec)、DeepSeek R1 Distill Llama 70B、Qwen2.5 72B、Llama 3.3 70B (Turbo, FP8)、Qwen3 30B (FP8)、NVIDIA Nemotron Nano 12B v2 VL (FP8)、Gemma 4 E4B (Non-reasoning)、NVIDIA Nemotron Nano 9B V2、Mistral Small 3.1、Llama Nemotron Super 49B v1.5 (Non-reasoning)、Qwen3 32B (Non-reasoning) (FP8)、Qwen3 32B (FP8)、Mistral Small 3.2 (FP8)、Llama 3.1 Nemotron 70B、Llama 3.1 8B (Turbo, FP8)、Llama 3.1 8B、Nemotron 3 Nano (Non-reasoning)、NVIDIA Nemotron Nano 9B V2 (Non-reasoning)、Qwen3 14B (Non-reasoning) (FP8)、Mistral Small 3、Qwen3 30B (Non-reasoning) (FP8)、Llama 3.1 70B、Llama 3.1 70B (Turbo, FP8)、Llama 4 Scout、Qwen3 14B (FP8)、Hermes 3 - Llama-3.1 70B、Phi-4、NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) (FP8)、Llama 3.2 11B (Vision)、Gemma 3 27B、Gemma 3 4B、Gemma 3 12B。
DeepInfraで利用できるモデルのうち、知能が最も高いのはIntelligence Indexスコア45のGLM-5.3 (max)です。
DeepInfraで出力速度が最も速いモデルは、毎秒243.7トークンのInkling Smallです。
DeepInfraで最初の回答トークンまでの時間が最短のモデルは、0.50秒のLlama 4 Maverick (FP8)です。遅延が短いほど、最初の応答が速くなります。
DeepInfraでブレンド料金が最も安いモデルは、100万トークンあたり$0.02のLlama 3.1 8B (Turbo, FP8)です(キャッシュヒット/入力/出力を7:2:1とした場合)。
DeepInfraのモデル間では料金に最大55倍の差があり、Llama 3.1 8B (Turbo, FP8)の100万トークンあたり$0.02から、Llama 3.1 Nemotron 70Bの$1.20までとなっています。
はい。DeepInfraはOpenAI互換APIを提供しているため、OpenAIからの切り替えや既存のOpenAI SDK連携の利用が容易です。
DeepInfraの108モデル中105モデルが、構造化出力のJSONモードに対応しています。
DeepInfraの108モデル中105モデルが関数呼び出し(ツール利用)に対応しています。
はい。DeepInfraは60推論モデルを提供しています:GLM-5.3 (max)、GLM-5.3-Flash、Qwen3.8 2.4T A95B、DeepSeek V4 Pro 0813 (max)、DeepSeek V4 Flash 0731 (max)、GLM-5.2 (max) (FP4)、Qwen3.8 27B (xhigh)、Kimi K2.6 (FP4)、DeepSeek V4 Pro (max) (FP4)、DeepSeek V4 Pro (high) (FP4)、MiniMax-M3、GLM-5 (FP4)、GLM-5.1 (FP4)、MiMo-V2.5-Pro、Kimi K2.7 Code、Inkling Small、Hy3 (FP8)、Inkling (FP8)、DeepSeek V4 Flash (high) (FP4)、DeepSeek V4 Flash (max) (FP4)、Kimi K2.5、Nemotron 3 Ultra、Nemotron 3 Ultra BF16、Qwen3.5 27B (FP8)、MiniMax-M2.5 (FP8)、MiMo-V2.5、GLM-4.7 (FP4)、Qwen3.6 27B FP8、DeepSeek V3.2 (FP4)、Ling 3.0 Flash、Step 3.7 Flash、Qwen3.5 35B A3B (FP8)、Qwen3.5 397B A17B (FP8)、Qwen3.6 35B A3B (FP8)、GLM-4.6 (FP4)、Muse Glimmer (high)、Gemma 4 26B A4B (FP8)、Qwen3.5 122B A10B (FP4)、Gemma 4 31B、GLM-4.7-Flash、Granite 4.2 30B、Nemotron 3 Super、Qwen3.5 4B (FP8)、DeepSeek R1 0528、Qwen3 235B A22B 2507 (FP8)、gpt-oss-120b (high) (Turbo)、gpt-oss-120b (high)、Granite 4.2 8B、gpt-oss-120b (low)、Granite 4.2 3B、gpt-oss-20b (high)、Llama Nemotron Super 49B v1.5、Gemma 4 E4B、Nemotron 3 Nano、DeepSeek R1 Distill Llama 70B、Qwen3 30B (FP8)、NVIDIA Nemotron Nano 12B v2 VL (FP8)、NVIDIA Nemotron Nano 9B V2、Qwen3 32B (FP8)、Qwen3 14B (FP8)。推論モデルは回答前に拡張思考を行い、複雑な問題に取り組みます。
はい。DeepInfraの全108モデルがオープンウェイトです。
はい。インフラストラクチャの変更、負荷分散、アップデートにより、プロバイダーの性能は時間とともに変化する場合があります。すべてのプロバイダーを継続的にベンチマークし、「推移」グラフに過去の性能傾向を表示しています。
DeepInfraのモデルを選ぶ際は、知能(品質を重視するタスク)、出力速度(高スループットが必要なタスク)、遅延(最初の応答の速さが必要な対話型アプリケーション)、料金(費用を重視するワークロード)、コンテキストウィンドウの規模、JSONモード、関数呼び出しへの対応などを検討してください。