Phi-4 Multimodal Instructの知能、性能、料金の分析
モデル概要
知能
速度
入力料金
出力料金
冗長性
Phi-4 Multimodal Instructの知能は平均を下回りますが、料金は割安です。これは同程度の規模の他の非推論オープンウェイトモデルと比較した場合です。 このモデルはテキスト、画像、音声入力に対応し、テキストを出力します。コンテキストウィンドウは128kトークンで、知識は2024年6月までです。
Phi-4 Multimodal InstructのArtificial Analysis Intelligence Indexスコアは5で、同等のモデルの中では平均以下です(中央値:6)。
Phi-4 Multimodal Instructの料金は入力100万トークンあたり$0.00(競争力が高い、中央値:$0.05)、出力100万トークンあたり$0.00(競争力が高い、中央値:$0.15)です。
Phi-4 Multimodal Instructの出力速度は毎秒18トークンで、特に遅い水準です(101)。
| 推論 | いいえ このページでは、このモデルの非推論版を表示しています。 推論版も存在する可能性があります。 |
|---|---|
| 入力モダリティ | 対応:テキスト、画像、音声 |
| 出力モダリティ | 対応:テキスト |
| 知識のカットオフ | 2024/06/01 |
| コンテキストウィンドウ | 128k Arial 12ポイントのA4用紙約192ページ分 |
| 総パラメーター数 | 5.6B |
| ライセンス | MIT |
| モデルウェイト | Hugging Face |
各指標は同じクラスのモデルと比較します。
- 非推論モデル → 他の非推論モデルとのみ比較
- 推論モデル → 推論モデルと非推論モデルの両方と比較
- オープンウェイトモデル → 同じ規模クラスの他のオープンウェイトモデルとのみ比較:
- 極小:パラメーター数≤4B
- 小:パラメーター数4B~40B
- 中:パラメーター数40B~150B
- 大:パラメーター数>150B
- 独自モデル → 入力/出力のブレンド料金を3:1として、同じ料金帯の独自モデルとオープンウェイトモデルを比較:
- 100万トークンあたり<$0.15
- 100万トークンあたり$0.15~$1
- 100万トークンあたり>$1
ハイライト
知能
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index by Open Weights / Proprietary
Intelligence Evaluations
Agentic real-world work tasks, (Elo-500)/2000
Agentic tool use
Agentic coding & terminal use
Coding
Reasoning & knowledge
Scientific reasoning
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Agentic knowledge work, Elo
Agentic SaaS workflows
Legal agentic work, criterion pass rate
Agentic business operations
Instruction following
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
Openness Index
Artificial Analysis Openness Index: Score
Intelligence Indexの比較
Intelligence Index vs. Cost per Intelligence Index Task
費用
Cost per Intelligence Index Task
Cost to Run Artificial Analysis Intelligence Index
Pricing: Cache Hit, Input, and Output
コンテキストウィンドウ
Context Window
速度
出力速度(1秒あたりのトークン数)で測定
Output Speed
Time per Intelligence Index Task
遅延
最初のトークンまでの時間(秒)で測定
Latency: Time To First Answer Token
エンドツーエンド応答時間
Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed
End-to-End Response Time
モデル規模(オープンウェイトモデルのみ)
Model Size: Total and Active Parameters
よくある質問
Phi-4 Multimodal Instructに関するよくある質問
Phi-4 Multimodal Instructは2025年2月26日にリリースされました。
Phi-4 Multimodal InstructはMicrosoftが開発しました。
Phi-4 Multimodal InstructのArtificial Analysis Intelligence Indexの推定スコアは5で、同程度の規模の他の非推論オープンウェイトモデルの中では平均以下水準です(中央値:6)。
Phi-4 Multimodal Instructの出力速度は毎秒17.7トークンです(MicrosoftのAPIに基づく)。同程度の規模の他の非推論オープンウェイトモデルと比べて下位水準です(中央値:100.7 t/s)。
Phi-4 Multimodal Instructの最初のトークンまでの時間(TTFT)は0.81秒です(MicrosoftのAPIに基づく)。同程度の規模の他の非推論オープンウェイトモデルと比べて非常に優れている水準です(中央値:1.72秒)。
いいえ。Phi-4 Multimodal Instructは推論モデルではありません。思考の連鎖による長い推論を行わず、直接回答します。
Phi-4 Multimodal Instructはテキスト、画像、音声入力に対応しています。
Phi-4 Multimodal Instructはテキスト出力に対応しています。
はい。Phi-4 Multimodal Instructは画像入力に対応し、画像の分析、説明、画像に関する質問への回答が可能です。
はい。Phi-4 Multimodal Instructはマルチモーダルです。テキスト、画像、音声入力を処理し、テキスト出力を生成できます。
Phi-4 Multimodal Instructのコンテキストウィンドウは130kトークンです。これは、モデルが1回のリクエストで処理できるテキストと会話履歴の量を決定します。
はい。Phi-4 Multimodal Instructはオープンウェイトモデルです。モデルウェイトが一般公開されており、ダウンロードしてセルフホストできます。
Phi-4 Multimodal Instructのパラメーター数は5.6Bです。
Phi-4 Multimodal InstructはMITライセンスで公開されており、商用利用が認められています。 ライセンスを表示
Phi-4 Multimodal InstructのArtificial Analysis Intelligence Indexスコアは5です。この複合ベンチマークでは、推論、知識、数学、コーディングについてモデルを評価します。
Phi-4 Multimodal Instructの知識のカットオフは2024年6月です。モデルの学習データには、この日付までの情報が含まれます。
はい。Phi-4 Multimodal Instructは1社のプロバイダーを通じてAPIで利用できます。 APIプロバイダーを比較
Phi-4 Multimodal Instructは1社のAPIプロバイダーを通じて利用できます。 プロバイダーを比較