DeepInfra:モデルの知能、性能、料金

この分析は、ユースケースに最適なDeepInfra提供モデルを選ぶための参考情報です。
最高の知能
UpdatedIntelligence Index
モデル合計:107件
最速
出力速度
モデル合計:107件
最安料金
ブレンド料金(100万トークンあたり)
モデル合計:107件
DeepInfraは107モデルを提供しており、知能、性能、料金の特性はそれぞれ異なります。 以下でモデル間の主要指標を比較します。
- DeepInfraで知能が上位のモデルはMiMo-V2.6-Pro(46)、GLM-5.3 (max)(45)、GLM-5.3-Flash(42)です。
- 出力速度が最も速いモデルはNemotron 3.5 Lightning (NVFP4)(295 t/s)、gpt-oss-120b (high) (Turbo)(240 t/s)、Granite 4.2 3B(220 t/s)です。 モデル間の速度差は大きく、最速と最遅で78%の差があります。
- 遅延では、Qwen3 30B (non-reasoning) (FP8)(0.57秒)、Qwen3.5 35B A3B (non-reasoning) FP8(0.58秒)、Qwen3 Coder 480B (Turbo, FP4)(0.67秒)の最初の回答トークンまでの時間が最短です。
- 料金では、Llama 3.1 8B (Turbo, FP8)($0.02)、Llama 3.1 8B($0.02)、Granite 4.2 3B($0.02)の100万トークンあたりのブレンド料金が最安です。
- DeepInfraで最大のコンテキストウィンドウに対応するモデルはGLM-5.2 (max) (FP4)(1M)、DeepSeek V4 Pro (max) (FP4)(1M)、DeepSeek V4 Flash (high) (FP4)(1M)です。
知能評価
Artificial Analysis Intelligence Index
Intelligence Evaluations
Agentic knowledge work, (Elo-500)/2000
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Agentic scientific research workflows in a terminal
Quantitative analysis on spreadsheets & documents
Kubernetes incident root-cause analysis
Visual reasoning
Medical long context reasoning
Intelligence Indexと料金
コンテキストウィンドウ
コンテキストウィンドウ
料金
Intelligence Indexと料金
性能の概要
出力速度と料金
速度
出力速度(1秒あたりのトークン数)で測定
出力速度
遅延
最初のトークンまでの時間(秒)で測定
遅延: 最初の回答トークンまでの時間
エンドツーエンド応答時間
Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed
エンドツーエンド応答時間と料金
詳細分析 | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
MiMo-V2.6-Pro | 1M | オープン | 46 | $0.12 | 26 | 1.72 | 99.53 | 78.24 | |||
GLM-5.3 (max) | 1.05M | オープン | 45 | $1.39 | 137 | 1.01 | 19.24 | 14.58 | |||
GLM-5.3-Flash | 1M | オープン | 42 | $0.34 | 24 | 1.46 | 105.74 | 83.42 | |||
Qwen3.8 2.4T A95B | 262k | オープン | 40 | $2.89 | 95 | 1.47 | 27.92 | 21.16 | |||
DeepSeek V4.1 Flash (max) | 1M | オープン | 39 | $0.44 | 79 | 0.85 | 32.54 | 25.35 | |||
MiMo-V2.6-Flash | 1M | オープン | 38 | $0.42 | 26 | 1.67 | 98.94 | 77.82 | |||
DeepSeek V4 Pro 0813 (max) | 1M | オープン | 36 | $1.00 | 56 | 1.27 | 45.78 | 35.61 | |||
DeepSeek V4 Flash Vision (max) | 1M | 独自 | 35 | $0.51 | 150 | 0.92 | 17.55 | 13.30 | |||
DeepSeek V4 Flash 0731 (max) | 1M | オープン | 34 | $0.10 | 42 | 1.04 | 59.93 | 47.11 | |||
GLM-5.2 (max) (FP4) | 1.05M | オープン | 34 | $0.37 | 90 | 1.12 | 28.93 | 22.25 | |||
Qwen3.8 27B (xhigh) | 262k | オープン | 34 | $0.52 | 29 | 1.54 | 87.35 | 68.65 | |||
DeepSeek V4 Pro (max) (FP4) | 1.05M | オープン | 30 | $1.19 | 37 | 1.37 | 132.72 | 117.88 | |||
DeepSeek V4 Pro (high) (FP4) | 65.5k | オープン | 30* | -- | 35 | 1.66 | 72.15 | 56.35 | |||
MiniMax-M3 | 524k | オープン | 29 | $0.44 | 25 | 0.91 | 102.44 | 81.22 | |||
GLM-5 (FP4) | 203k | オープン | 28* | -- | 78 | 1.28 | 47.68 | 39.96 | |||
Kimi K2.6 (FP4) | 262k | オープン | 27 | $0.51 | 20 | 1.67 | 243.21 | 217.15 | |||
GLM-5.1 (FP4) | 203k | オープン | 26 | $0.58 | 41 | 1.32 | 106.17 | 92.64 | |||
MiMo-V2.5-Pro | 65.5k | オープン | 26 | $0.28 | 26 | 2.09 | 97.60 | 76.40 | |||
Kimi K2.7 Code | 262k | オープン | 26 | $0.39 | 25 | 1.08 | 108.25 | 87.54 | |||
Inkling Small | 524k | オープン | 26 | -- | 238 | 0.60 | 11.11 | 8.41 | |||
Hy3 (FP8) | 262k | オープン | 25 | $0.07 | 33 | 1.03 | 76.02 | 59.99 | |||
MiMo-V2.5 | 262k | オープン | 25* | -- | 28 | 1.57 | 89.64 | 70.46 | |||
Inkling (xhigh) (FP8) | 131k | オープン | 25 | -- | 128 | 0.76 | 20.23 | 15.57 | |||
DeepSeek V4 Flash (high) (FP4) | 1.05M | オープン | 24 | $0.08 | 35 | 1.54 | 50.66 | 35.01 | |||
GLM-5.1 (non-reasoning) (FP4) | 203k | オープン | 24* | -- | 48 | 1.19 | 11.54 | -- | |||
DeepSeek V4 Flash (max) (FP4) | 1.05M | オープン | 24 | $0.07 | 36 | 1.37 | 173.39 | 157.95 | |||
Kimi K2.6 (non-reasoning) (FP4) | 262k | オープン | 24* | -- | 18 | 1.54 | 29.23 | -- | |||
Kimi K2.5 | 262k | オープン | 23* | -- | 25 | 1.35 | 140.24 | 118.86 | |||
Nemotron 3 Ultra | 262k | オープン | 23 | $0.47 | 66 | 1.27 | 43.10 | 34.30 | |||
Nemotron 3 Ultra BF16 | 262k | オープン | 23 | $0.99 | 85 | 1.68 | 34.48 | 26.89 | |||
Qwen3.5 27B (FP8) | 262k | オープン | 23* | -- | 61 | 1.01 | 42.00 | 32.79 | |||
MiniMax-M2.5 (FP8) | 197k | オープン | 23* | -- | 27 | 0.91 | 94.98 | 75.26 | |||
GLM-4.7 (FP4) | 203k | オープン | 22* | -- | 17 | 1.33 | 150.59 | 119.41 | |||
GLM-5 (non-reasoning) (FP8) | 203k | オープン | 22* | -- | 71 | 1.09 | 8.12 | -- | |||
DeepSeek V3.2 (FP4) | 164k | オープン | 21* | -- | 16 | 1.53 | 156.57 | 124.03 | |||
Qwen3.5 397B A17B (non-reasoning) (FP8) | 262k | オープン | 21* | -- | 35 | 1.15 | 15.44 | -- | |||
Qwen3.6 27B FP8 | 262k | オープン | 21 | $0.37 | 63 | 1.03 | 98.44 | 89.52 | |||
Ling 3.0 Flash | 131k | オープン | 20 | $0.02 | 48 | 0.75 | 52.98 | 41.78 | |||
Qwen3.6 27B (non-reasoning) FP8 | 262k | オープン | 20* | -- | 61 | 1.01 | 9.17 | -- | |||
Step 3.7 Flash | 256k | オープン | 19* | -- | 169 | 0.71 | 15.47 | 11.81 | |||
Qwen3.5 27B (non-reasoning) FP8 | 262k | オープン | 19* | -- | 60 | 1.00 | 9.27 | -- | |||
Qwen3.5 35B A3B (FP8) | 262k | オープン | 19* | -- | 120 | 0.67 | 21.47 | 16.64 | |||
GLM-4.6 (FP4) | 203k | オープン | 19* | -- | 19 | 1.71 | 130.18 | 102.77 | |||
Qwen3.5 397B A17B (FP8) | 262k | オープン | 18 | -- | 38 | 1.22 | 98.11 | 83.74 | |||
MiMo-V2.5-Pro (non-reasoning) | 65.5k | オープン | 18* | -- | 29 | 1.47 | 18.58 | -- | |||
Qwen3.6 35B A3B (FP8) | 262k | オープン | 18 | $0.14 | 32 | 0.85 | 183.26 | 166.94 | |||
Qwen3.5 122B A10B (non-reasoning) (FP4) | 262k | オープン | 18* | -- | 148 | 0.70 | 4.07 | -- | |||
Muse Glimmer (high) | 131k | オープン | 17 | $0.05 | 143 | 0.86 | 18.33 | 13.98 | |||
GLM-4.7 (non-reasoning) (FP4) | 203k | オープン | 17* | -- | 19 | 1.49 | 27.60 | -- | |||
Gemma 4 26B A4B (FP8) | 262k | オープン | 17* | -- | 35 | 0.80 | 73.17 | 57.89 | |||
DeepSeek V3.2 (non-reasoning) | 164k | オープン | 16* | -- | 14 | 1.62 | 38.33 | -- | |||
Qwen3.5 122B A10B (FP4) | 262k | オープン | 16 | $0.24 | 163 | 0.68 | 15.99 | 12.25 | |||
Qwen3.6 35B A3B (non-reasoning) (FP8) | 262k | オープン | 15* | -- | 28 | 1.00 | 18.87 | -- | |||
Qwen3.5 35B A3B (non-reasoning) FP8 | 262k | オープン | 15* | -- | 139 | 0.70 | 4.28 | -- | |||
GLM-4.7-Flash | 203k | オープン | 15* | -- | 20 | 1.32 | 127.66 | 101.07 | |||
Granite 4.2 30B | 131k | オープン | 15* | -- | 77 | 0.88 | 33.27 | 25.92 | |||
Gemma 4 31B | 262k | オープン | 15 | $0.10 | 15 | 2.72 | 148.94 | 113.52 | |||
DeepSeek V3.1 Terminus (non-reasoning) (FP4) | 164k | オープン | 14* | -- | 40 | 1.17 | 13.58 | -- | |||
Gemma 4 31B (non-reasoning) (FP8) | 262k | オープン | 14* | -- | 15 | 2.23 | 35.13 | -- | |||
DeepSeek V3.1 (non-reasoning) (FP4) | 164k | オープン | 14* | -- | 9 | 1.46 | 59.76 | -- | |||
Gemma 4 26B A4B (non-reasoning) (FP8) | 262k | オープン | 13* | -- | 31 | 0.72 | 16.91 | -- | |||
Qwen3.5 4B (FP8) | 262k | オープン | 13* | -- | 31 | 0.87 | 82.79 | 65.53 | |||
DeepSeek R1 0528 | 164k | オープン | 13* | -- | 31 | 0.89 | 82.75 | 65.49 | |||
Nemotron 3.5 Lightning (NVFP4) | 262k | オープン | 13 | $0.12 | 337 | 0.52 | 7.94 | 5.93 | |||
Nemotron 3 Super | 262k | オープン | 13 | $0.48 | 69 | 77.50 | 113.81 | 29.04 | |||
Qwen3 235B A22B 2507 (FP8) | 262k | オープン | 13 | -- | 35 | 0.82 | 72.18 | 57.09 | |||
Qwen3 235B 2507 (FP8) | 262k | オープン | 12* | -- | 15 | 1.37 | 35.59 | -- | |||
Qwen3 Coder 480B (Turbo, FP4) | 262k | オープン | 12* | -- | 39 | 0.94 | 13.70 | -- | |||
gpt-oss-120b (high) (Turbo) | 131k | オープン | 12 | $0.11 | 322 | 0.73 | 8.48 | 6.20 | |||
gpt-oss-120b (high) | 131k | オープン | 12 | $0.03 | 52 | 0.60 | 48.97 | 38.70 | |||
Granite 4.2 8B | 131k | オープン | 11 | $0.02 | 65 | 0.73 | 39.01 | 30.62 | |||
Qwen3.5 4B (non-reasoning) FP8 | 262k | オープン | 11* | -- | 16 | 0.94 | 31.67 | -- | |||
gpt-oss-120b (low) | 131k | オープン | 10* | -- | 44 | 0.72 | 56.96 | 44.99 | |||
Llama 4 Maverick (FP8) | 1.05M | オープン | 10* | -- | 68 | 0.58 | 7.99 | -- | |||
DeepSeek V3 0324 (FP4) | 164k | オープン | 10 | $0.02 | 48 | 1.12 | 11.52 | -- | |||
Qwen3 Next 80B A3B | 262k | オープン | 10* | -- | 147 | 0.67 | 4.08 | -- | |||
Granite 4.2 3B | 131k | オープン | 9 | $0.01 | 223 | 0.52 | 11.71 | 8.95 | |||
gpt-oss-20b (high) | 131k | オープン | 9 | $0.01 | 111 | 0.50 | 23.09 | 18.07 | |||
Gemma 4 E4B | 262k | オープン | 9* | -- | 70 | 0.84 | 36.80 | 28.77 | |||
Nemotron 3 Nano | 262k | オープン | 9 | $0.02 | 105 | 6.23 | 30.00 | 19.02 | |||
Qwen3 32B (FP8) | 41k | オープン | 9* | -- | 31 | 1.50 | 81.16 | 63.73 | |||
DeepSeek V3 (Dec) | 164k | オープン | 8 | $0.02 | 19 | 2.12 | 27.88 | -- | |||
Mistral Small 3.2 (FP8) | 128k | オープン | 8* | -- | 37 | 0.96 | 14.41 | -- | |||
Qwen3 14B (FP8) | 32.8k | オープン | 8* | -- | 27 | 1.56 | 92.73 | 72.93 | |||
Llama 4 Scout | 328k | オープン | 8* | -- | 33 | 0.83 | 15.82 | -- | |||
DeepSeek R1 Distill Llama 70B | 131k | オープン | 8* | -- | -- | -- | -- | -- | |||
Qwen2.5 72B | 32.8k | オープン | 8* | -- | 26 | 2.48 | 21.42 | -- | |||
Llama 3.3 70B (Turbo, FP8) | 131k | オープン | 8* | -- | 17 | 2.26 | 31.79 | -- | |||
Qwen3 30B (FP8) | 41k | オープン | 8* | -- | 126 | 0.58 | 20.36 | 15.83 | |||
Gemma 4 E4B (non-reasoning) | 262k | オープン | 7* | -- | 72 | 0.79 | 7.73 | -- | |||
NVIDIA Nemotron Nano 9B V2 | 131k | オープン | 7* | -- | 104 | 5.12 | 29.19 | 19.26 | |||
Qwen3 32B (non-reasoning) (FP8) | 41k | オープン | 7* | -- | 30 | 1.34 | 17.90 | -- | |||
Mistral Small 3.1 | 128k | オープン | 7 | $0.02 | 39 | 0.95 | 13.82 | -- | |||
Llama 3.1 8B (Turbo, FP8) | 131k | オープン | 7* | -- | 18 | 1.16 | 28.67 | -- | |||
Llama 3.1 8B | 131k | オープン | 7* | -- | 19 | 1.13 | 27.39 | -- | |||
Nemotron 3 Nano (non-reasoning) | 262k | オープン | 7* | -- | 105 | 0.69 | 5.44 | -- | |||
NVIDIA Nemotron Nano 9B V2 (non-reasoning) | 131k | オープン | 7* | -- | 99 | 6.30 | 11.33 | -- | |||
Qwen3 14B (non-reasoning) (FP8) | 41k | オープン | 7* | -- | 29 | 1.21 | 18.45 | -- | |||
Mistral Small 3 | 32.8k | オープン | 7* | -- | 51 | 0.96 | 10.67 | -- | |||
Qwen3 30B (non-reasoning) (FP8) | 41k | オープン | 7* | -- | 148 | 0.58 | 3.95 | -- | |||
Llama 3.1 70B | 131k | オープン | 7* | -- | 40 | 1.96 | 14.39 | -- | |||
Llama 3.1 70B (Turbo, FP8) | 131k | オープン | 7* | -- | 39 | 2.02 | 14.97 | -- | |||
Hermes 3 - Llama-3.1 70B | 131k | オープン | 6* | -- | 28 | 2.28 | 19.88 | -- | |||
Phi-4 | 16.4k | オープン | 6* | -- | 70 | 0.96 | 8.06 | -- | |||
Llama 3.2 11B (Vision) | 131k | オープン | 5* | -- | 15 | 3.03 | 35.78 | -- | |||
Gemma 3 27B | 131k | オープン | 5 | $0.16 | 25 | 1.25 | 21.45 | -- | |||
Gemma 3 4B | 131k | オープン | 5* | -- | 21 | 1.55 | 25.74 | -- | |||
Llama 3 8B | 8.19k | オープン | 5* | -- | -- | -- | -- | -- | |||
Gemma 3 12B | 131k | オープン | 4 | $0.13 | 43 | 1.09 | 12.77 | -- | |||
主要な定義
よくある質問
DeepInfraに関するよくある質問
DeepInfraが提供し、当社が追跡しているモデルは107モデルです:MiMo-V2.6-Pro、GLM-5.3 (max)、GLM-5.3-Flash、Qwen3.8 2.4T A95B、DeepSeek V4.1 Flash (max)、MiMo-V2.6-Flash、DeepSeek V4 Pro 0813 (max)、DeepSeek V4 Flash Vision (max)、DeepSeek V4 Flash 0731 (max)、GLM-5.2 (max) (FP4)、Qwen3.8 27B (xhigh)、DeepSeek V4 Pro (max) (FP4)、DeepSeek V4 Pro (high) (FP4)、MiniMax-M3、GLM-5 (FP4)、Kimi K2.6 (FP4)、GLM-5.1 (FP4)、MiMo-V2.5-Pro、Kimi K2.7 Code、Inkling Small、Hy3 (FP8)、MiMo-V2.5、Inkling (xhigh) (FP8)、DeepSeek V4 Flash (high) (FP4)、GLM-5.1 (non-reasoning) (FP4)、DeepSeek V4 Flash (max) (FP4)、Kimi K2.6 (non-reasoning) (FP4)、Kimi K2.5、Nemotron 3 Ultra、Nemotron 3 Ultra BF16、Qwen3.5 27B (FP8)、MiniMax-M2.5 (FP8)、GLM-4.7 (FP4)、GLM-5 (non-reasoning) (FP8)、DeepSeek V3.2 (FP4)、Qwen3.5 397B A17B (non-reasoning) (FP8)、Qwen3.6 27B FP8、Ling 3.0 Flash、Qwen3.6 27B (non-reasoning) FP8、Step 3.7 Flash、Qwen3.5 27B (non-reasoning) FP8、Qwen3.5 35B A3B (FP8)、GLM-4.6 (FP4)、Qwen3.5 397B A17B (FP8)、MiMo-V2.5-Pro (non-reasoning)、Qwen3.6 35B A3B (FP8)、Qwen3.5 122B A10B (non-reasoning) (FP4)、Muse Glimmer (high)、GLM-4.7 (non-reasoning) (FP4)、Gemma 4 26B A4B (FP8)、DeepSeek V3.2 (non-reasoning)、Qwen3.5 122B A10B (FP4)、Qwen3.6 35B A3B (non-reasoning) (FP8)、Qwen3.5 35B A3B (non-reasoning) FP8、GLM-4.7-Flash、Granite 4.2 30B、Gemma 4 31B、DeepSeek V3.1 Terminus (non-reasoning) (FP4)、Gemma 4 31B (non-reasoning) (FP8)、DeepSeek V3.1 (non-reasoning) (FP4)、Gemma 4 26B A4B (non-reasoning) (FP8)、Qwen3.5 4B (FP8)、DeepSeek R1 0528、Nemotron 3.5 Lightning (NVFP4)、Nemotron 3 Super、Qwen3 235B A22B 2507 (FP8)、Qwen3 235B 2507 (FP8)、Qwen3 Coder 480B (Turbo, FP4)、gpt-oss-120b (high) (Turbo)、gpt-oss-120b (high)、Granite 4.2 8B、Qwen3.5 4B (non-reasoning) FP8、gpt-oss-120b (low)、Llama 4 Maverick (FP8)、DeepSeek V3 0324 (FP4)、Qwen3 Next 80B A3B、Granite 4.2 3B、gpt-oss-20b (high)、Gemma 4 E4B、Nemotron 3 Nano、Qwen3 32B (FP8)、DeepSeek V3 (Dec)、Mistral Small 3.2 (FP8)、Qwen3 14B (FP8)、Llama 4 Scout、Qwen2.5 72B、Llama 3.3 70B (Turbo, FP8)、Qwen3 30B (FP8)、Gemma 4 E4B (non-reasoning)、NVIDIA Nemotron Nano 9B V2、Qwen3 32B (non-reasoning) (FP8)、Mistral Small 3.1、Llama 3.1 8B (Turbo, FP8)、Llama 3.1 8B、Nemotron 3 Nano (non-reasoning)、NVIDIA Nemotron Nano 9B V2 (non-reasoning)、Qwen3 14B (non-reasoning) (FP8)、Mistral Small 3、Qwen3 30B (non-reasoning) (FP8)、Llama 3.1 70B、Llama 3.1 70B (Turbo, FP8)、Hermes 3 - Llama-3.1 70B、Phi-4、Llama 3.2 11B (Vision)、Gemma 3 27B、Gemma 3 4B、Gemma 3 12B。
DeepInfraで利用できるモデルのうち、知能が最も高いのはIntelligence Indexスコア46のMiMo-V2.6-Proです。
DeepInfraで出力速度が最も速いモデルは、毎秒294.6トークンのNemotron 3.5 Lightning (NVFP4)です。
DeepInfraで最初の回答トークンまでの時間が最短のモデルは、0.57秒のQwen3 30B (non-reasoning) (FP8)です。遅延が短いほど、最初の応答が速くなります。
DeepInfraでブレンド料金が最も安いモデルは、100万トークンあたり$0.02のLlama 3.1 8B (Turbo, FP8)です(キャッシュヒット/入力/出力を7:2:1とした場合)。
DeepInfraのモデル間では料金に最大52倍の差があり、Llama 3.1 8B (Turbo, FP8)の100万トークンあたり$0.02から、Qwen3.8 2.4T A95Bの$1.14までとなっています。
はい。DeepInfraはOpenAI互換APIを提供しているため、OpenAIからの切り替えや既存のOpenAI SDK連携の利用が容易です。
はい。DeepInfraの全107モデルが、構造化出力のJSONモードに対応しています。
DeepInfraの107モデル中104モデルが関数呼び出し(ツール利用)に対応しています。
はい。DeepInfraは62推論モデルを提供しています:MiMo-V2.6-Pro、GLM-5.3 (max)、GLM-5.3-Flash、Qwen3.8 2.4T A95B、DeepSeek V4.1 Flash (max)、MiMo-V2.6-Flash、DeepSeek V4 Pro 0813 (max)、DeepSeek V4 Flash Vision (max)、DeepSeek V4 Flash 0731 (max)、GLM-5.2 (max) (FP4)、Qwen3.8 27B (xhigh)、DeepSeek V4 Pro (max) (FP4)、DeepSeek V4 Pro (high) (FP4)、MiniMax-M3、GLM-5 (FP4)、Kimi K2.6 (FP4)、GLM-5.1 (FP4)、MiMo-V2.5-Pro、Kimi K2.7 Code、Inkling Small、Hy3 (FP8)、MiMo-V2.5、Inkling (xhigh) (FP8)、DeepSeek V4 Flash (high) (FP4)、DeepSeek V4 Flash (max) (FP4)、Kimi K2.5、Nemotron 3 Ultra、Nemotron 3 Ultra BF16、Qwen3.5 27B (FP8)、MiniMax-M2.5 (FP8)、GLM-4.7 (FP4)、DeepSeek V3.2 (FP4)、Qwen3.6 27B FP8、Ling 3.0 Flash、Step 3.7 Flash、Qwen3.5 35B A3B (FP8)、GLM-4.6 (FP4)、Qwen3.5 397B A17B (FP8)、Qwen3.6 35B A3B (FP8)、Muse Glimmer (high)、Gemma 4 26B A4B (FP8)、Qwen3.5 122B A10B (FP4)、GLM-4.7-Flash、Granite 4.2 30B、Gemma 4 31B、Qwen3.5 4B (FP8)、DeepSeek R1 0528、Nemotron 3.5 Lightning (NVFP4)、Nemotron 3 Super、Qwen3 235B A22B 2507 (FP8)、gpt-oss-120b (high) (Turbo)、gpt-oss-120b (high)、Granite 4.2 8B、gpt-oss-120b (low)、Granite 4.2 3B、gpt-oss-20b (high)、Gemma 4 E4B、Nemotron 3 Nano、Qwen3 32B (FP8)、Qwen3 14B (FP8)、Qwen3 30B (FP8)、NVIDIA Nemotron Nano 9B V2。推論モデルは回答前に拡張思考を行い、複雑な問題に取り組みます。
はい。DeepInfraの107モデル中106モデルがオープンウェイトです:MiMo-V2.6-Pro、GLM-5.3 (max)、GLM-5.3-Flash、Qwen3.8 2.4T A95B、DeepSeek V4.1 Flash (max)、MiMo-V2.6-Flash、DeepSeek V4 Pro 0813 (max)、DeepSeek V4 Flash 0731 (max)、GLM-5.2 (max) (FP4)、Qwen3.8 27B (xhigh)、DeepSeek V4 Pro (max) (FP4)、DeepSeek V4 Pro (high) (FP4)、MiniMax-M3、GLM-5 (FP4)、Kimi K2.6 (FP4)、GLM-5.1 (FP4)、MiMo-V2.5-Pro、Kimi K2.7 Code、Inkling Small、Hy3 (FP8)、MiMo-V2.5、Inkling (xhigh) (FP8)、DeepSeek V4 Flash (high) (FP4)、GLM-5.1 (non-reasoning) (FP4)、DeepSeek V4 Flash (max) (FP4)、Kimi K2.6 (non-reasoning) (FP4)、Kimi K2.5、Nemotron 3 Ultra、Nemotron 3 Ultra BF16、Qwen3.5 27B (FP8)、MiniMax-M2.5 (FP8)、GLM-4.7 (FP4)、GLM-5 (non-reasoning) (FP8)、DeepSeek V3.2 (FP4)、Qwen3.5 397B A17B (non-reasoning) (FP8)、Qwen3.6 27B FP8、Ling 3.0 Flash、Qwen3.6 27B (non-reasoning) FP8、Step 3.7 Flash、Qwen3.5 27B (non-reasoning) FP8、Qwen3.5 35B A3B (FP8)、GLM-4.6 (FP4)、Qwen3.5 397B A17B (FP8)、MiMo-V2.5-Pro (non-reasoning)、Qwen3.6 35B A3B (FP8)、Qwen3.5 122B A10B (non-reasoning) (FP4)、Muse Glimmer (high)、GLM-4.7 (non-reasoning) (FP4)、Gemma 4 26B A4B (FP8)、DeepSeek V3.2 (non-reasoning)、Qwen3.5 122B A10B (FP4)、Qwen3.6 35B A3B (non-reasoning) (FP8)、Qwen3.5 35B A3B (non-reasoning) FP8、GLM-4.7-Flash、Granite 4.2 30B、Gemma 4 31B、DeepSeek V3.1 Terminus (non-reasoning) (FP4)、Gemma 4 31B (non-reasoning) (FP8)、DeepSeek V3.1 (non-reasoning) (FP4)、Gemma 4 26B A4B (non-reasoning) (FP8)、Qwen3.5 4B (FP8)、DeepSeek R1 0528、Nemotron 3.5 Lightning (NVFP4)、Nemotron 3 Super、Qwen3 235B A22B 2507 (FP8)、Qwen3 235B 2507 (FP8)、Qwen3 Coder 480B (Turbo, FP4)、gpt-oss-120b (high) (Turbo)、gpt-oss-120b (high)、Granite 4.2 8B、Qwen3.5 4B (non-reasoning) FP8、gpt-oss-120b (low)、Llama 4 Maverick (FP8)、DeepSeek V3 0324 (FP4)、Qwen3 Next 80B A3B、Granite 4.2 3B、gpt-oss-20b (high)、Gemma 4 E4B、Nemotron 3 Nano、Qwen3 32B (FP8)、DeepSeek V3 (Dec)、Mistral Small 3.2 (FP8)、Qwen3 14B (FP8)、Llama 4 Scout、Qwen2.5 72B、Llama 3.3 70B (Turbo, FP8)、Qwen3 30B (FP8)、Gemma 4 E4B (non-reasoning)、NVIDIA Nemotron Nano 9B V2、Qwen3 32B (non-reasoning) (FP8)、Mistral Small 3.1、Llama 3.1 8B (Turbo, FP8)、Llama 3.1 8B、Nemotron 3 Nano (non-reasoning)、NVIDIA Nemotron Nano 9B V2 (non-reasoning)、Qwen3 14B (non-reasoning) (FP8)、Mistral Small 3、Qwen3 30B (non-reasoning) (FP8)、Llama 3.1 70B、Llama 3.1 70B (Turbo, FP8)、Hermes 3 - Llama-3.1 70B、Phi-4、Llama 3.2 11B (Vision)、Gemma 3 27B、Gemma 3 4B、Gemma 3 12B。
はい。インフラストラクチャの変更、負荷分散、アップデートにより、プロバイダーの性能は時間とともに変化する場合があります。すべてのプロバイダーを継続的にベンチマークし、「推移」グラフに過去の性能傾向を表示しています。
DeepInfraのモデルを選ぶ際は、知能(品質を重視するタスク)、出力速度(高スループットが必要なタスク)、遅延(最初の応答の速さが必要な対話型アプリケーション)、料金(費用を重視するワークロード)、コンテキストウィンドウの規模、JSONモード、関数呼び出しへの対応などを検討してください。