PrimaLabs:モデルの知能、性能、料金
この分析は、ユースケースに最適なPrimaLabs提供モデルを選ぶための参考情報です。
最高の知能
UpdatedIntelligence Index
モデル合計:1件
最速
出力速度
モデル合計:1件
最安料金
ブレンド料金(100万トークンあたり)
モデル合計:1件
現在、PrimaLabsはDeepSeek V4 Flash 0731 (max)を提供しています。
ハイライト
知能評価
Artificial Analysis Intelligence Index
Intelligence Evaluations
Agentic knowledge work, (Elo-500)/2000
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Quantitative analysis on spreadsheets & documents
Agentic tool use
Kubernetes incident root-cause analysis
Visual reasoning
Medical long context reasoning
Intelligence Index vs. Price
コンテキストウィンドウ
Context Window
料金
Intelligence Index vs. Price
性能の概要
Output Speed vs. Price
速度
出力速度(1秒あたりのトークン数)で測定
Output Speed
遅延
最初のトークンまでの時間(秒)で測定
Latency: Time To First Answer Token
エンドツーエンド応答時間
Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed
End-to-End Response Time vs. Price
主要な定義
よくある質問
PrimaLabsに関するよくある質問
PrimaLabsが提供し、当社が追跡しているモデルは1モデルです:DeepSeek V4 Flash 0731 (max)。
PrimaLabsで利用できるモデルのうち、知能が最も高いのはIntelligence Indexスコア35のDeepSeek V4 Flash 0731 (max)です。
PrimaLabsで出力速度が最も速いモデルは、毎秒341.7トークンのDeepSeek V4 Flash 0731 (max)です。
PrimaLabsで最初の回答トークンまでの時間が最短のモデルは、6.78秒のDeepSeek V4 Flash 0731 (max)です。遅延が短いほど、最初の応答が速くなります。
PrimaLabsでブレンド料金が最も安いモデルは、100万トークンあたり$0.06のDeepSeek V4 Flash 0731 (max)です(キャッシュヒット/入力/出力を7:2:1とした場合)。
はい。PrimaLabsはOpenAI互換APIを提供しているため、OpenAIからの切り替えや既存のOpenAI SDK連携の利用が容易です。
はい。PrimaLabsの全1モデルが、構造化出力のJSONモードに対応しています。
はい。PrimaLabsの全1モデルが関数呼び出し(ツール利用)に対応しています。
はい。PrimaLabsは1推論モデルを提供しています:DeepSeek V4 Flash 0731 (max)。推論モデルは回答前に拡張思考を行い、複雑な問題に取り組みます。
はい。PrimaLabsの全1モデルがオープンウェイトです。
はい。インフラストラクチャの変更、負荷分散、アップデートにより、プロバイダーの性能は時間とともに変化する場合があります。すべてのプロバイダーを継続的にベンチマークし、「推移」グラフに過去の性能傾向を表示しています。
PrimaLabsのモデルを選ぶ際は、知能(品質を重視するタスク)、出力速度(高スループットが必要なタスク)、遅延(最初の応答の速さが必要な対話型アプリケーション)、料金(費用を重視するワークロード)、コンテキストウィンドウの規模、JSONモード、関数呼び出しへの対応などを検討してください。