モデル比較:知能、性能、料金の分析
試す知能
Updated知能Updated
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index:オープンウェイトとプロプライエタリ
特定の能力や業界におけるモデルの性能を測定
Artificial Analysis 財務・会計指数
Intelligence Evaluations
Agentic knowledge work, (Elo-500)/2000
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Agentic scientific research workflows in a terminal
Quantitative analysis on spreadsheets & documents
Kubernetes incident root-cause analysis
Visual reasoning
Medical long context reasoning
AA-Briefcase v1.1Updated
AA-Briefcase Elo
AA-Omniscience
AA-Omniscience Index
Openness Index
Artificial Analysis Openness Index:スコア
Intelligence Indexの比較
Intelligence Index とタスクあたりのコスト
トークン使用量
Intelligence Indexのタスクあたりの出力トークン数
費用
Intelligence Index タスクあたりのコスト
Artificial Analysis Intelligence Index の実行コスト
料金:キャッシュヒット・入力・出力
コンテキストウィンドウ
コンテキストウィンドウ
速度
出力速度(1秒あたりのトークン数)で測定
出力速度
Intelligence Indexのタスクあたりの時間
遅延
最初のトークンまでの時間(秒)で測定
遅延: 最初の回答トークンまでの時間
エンドツーエンド応答時間
Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed
エンドツーエンド応答時間
モデル規模(オープンウェイトモデルのみ)
モデルサイズ:総パラメータ数とアクティブパラメータ数
よくある質問
現在、Claude Opus 5.5 (Max, Default Fallback)が58のスコアでArtificial Analysis Intelligence Indexの首位です。評価対象は198モデルです。
Intelligence Indexで上位のAIモデルは1位:Claude Opus 5.5 (Max, Default Fallback)(58)、2位:Claude Sonnet 5.5 (Max, Default Fallback)(56)、3位:Claude Opus 5.5 (Xhigh, Default Fallback)(56)、4位:Claude Opus 5.5 (High, Default Fallback)(54)、5位:Claude Fable 5.1 (Max, Default Fallback)(53)です。
Celeris-1が毎秒1,656.6トークンで最速で、Mercury 2(698.1 t/s)、Mercury 2.5(660.6 t/s)が続きます。
Llama 3.1 Instruct 8Bが100万トークンあたり$0.02(ブレンド料金)で最も安価で、Granite 4.2 3B($0.02)、Nova Micro($0.03)が続きます。
Gemini 2.5 Flash-Lite (Non-reasoning)の最初のトークンまでの時間が0.31秒で最短で、Nemotron 3 Nano Omni 30B A3B Reasoning(0.35秒)、North Mini Code(0.42秒)が続きます。
MiMo-V2.6-ProがIntelligence Indexスコア46で、オープンウェイトモデルの最上位です。評価対象の全198モデル中、オープンウェイトモデルは81モデルです。
Intelligence Indexで上位のオープンウェイトAIモデルは1位:MiMo-V2.6-Pro(46)、2位:GLM-5.3 (Max)(45)、3位:Kimi K3 (Max)(44)です。
Claude Opus 5.5 (Max, Default Fallback)がIntelligence Indexスコア58で、176推論モデルの首位です。推論モデルは回答前に拡張思考を行い、複雑な問題に取り組みます。
知能(品質)、料金、出力速度(1秒あたりのトークン数)、遅延(最初のトークンまでの時間)、エンドツーエンド応答時間、コンテキストウィンドウの規模など、複数の側面でモデルを比較します。性能指標は、標準化したプロンプトを使って690モデルで直接測定します。
グラフ内のモデル名または行をクリックすると、詳細な指標と類似モデルとの直接比較を掲載した専用ページを表示できます。モデル選択機能を使って、各グラフに表示するモデルを変更することもできます。 リーダーボードを表示