| 知能 | |||
|---|---|---|---|
| Intelligence Index | 45 | 53 | |
| AA-Briefcase v1.1 | 1621 | 1676 | |
| GDPval-AA v2.1 | 1671 | 1758 | |
| AutomationBench-AA | 56% | 59% | |
| Terminal-Bench 4.0 | 39% | 52% | |
| SciCode | 52% | 63% | |
| Humanity's Last Exam | 43% | 59% | |
| GDP.pdf | 23% | 26% | |
| CritPt | 18% | 30% | |
| AA-Omniscience | 12 | 43 | |
| AA-LCR v1.1 | 80% | 85% | |
| 費用 | |||
| 100万トークンあたりの料金 | $1.175 | $7.175 | |
| 入力100万トークンあたりの価格 | $2.00 | $10.00 | |
| 出力100万トークンあたりの価格 | $6.00 | $50.00 | |
| キャッシュヒット100万トークンあたりの価格 | $0.25 | $0.25 | |
| タスクあたりのコスト | $5.41 | $7.63 | |
| Intelligence Index 実行コスト | $4,935 | $13,129 | |
| トークン使用量 | |||
| タスクあたりの出力トークン | 108k | 78k | |
| タスクあたりの推論トークン | 71k | 47k | |
| Intelligence Index 実行の出力トークン | 192M | 188M | |
| パフォーマンス | |||
| 出力速度 | 38トークン/秒 | 62トークン/秒 | |
| 最初のトークンまでの時間 | 2.69秒 | 259.51秒 | |
| 最初の回答トークンまでの時間 | 55.94秒 | 259.51秒 | |
| エンドツーエンド応答時間 | 69.26秒 | 267.51秒 | |
| タスクあたりの時間 | 2343.76秒 | 768.35秒 | |
| 技術仕様 | |||
| コンテキストウィンドウ | 984kトークンArial 12ポイントのA4用紙約1,475ページ分 | 1000kトークンArial 12ポイントのA4用紙約1,500ページ分 | |
| リリース日 | 2026年9月 | 2026年9月 | |
| 推論 | はい | はい | |
| 入力モダリティ | 対応:テキスト、画像 | 対応:テキスト、画像 | |
| 出力モダリティ | 対応:テキスト | 対応:テキスト | |
| オープンウェイト | いいえ | いいえ | |
知能Updated
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index:オープンウェイトとプロプライエタリ
特定の能力や業界におけるモデルの性能を測定
Artificial Analysis 財務・会計指数
Intelligence Evaluations
Agentic knowledge work, (Elo-500)/2000
Agentic real-world work tasks, (Elo-500)/2000
Agentic SaaS workflows
Agentic coding & terminal use
Coding
Reasoning & knowledge
Professional document reasoning, All-pass
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Legal agentic work, criterion pass rate
Agentic business operations
Agentic scientific research workflows in a terminal
Quantitative analysis on spreadsheets & documents
Kubernetes incident root-cause analysis
Visual reasoning
Medical long context reasoning
AA-Briefcase v1.1Updated
AA-Briefcase Elo
AA-Omniscience
AA-Omniscience Index
Intelligence Indexの比較
Intelligence Index とタスクあたりのコスト
タスク当たりのコスト(USD、対数スケール)
トークン使用量
Intelligence Indexのタスクあたりの出力トークン数
費用
Intelligence Index タスクあたりのコスト
Artificial Analysis Intelligence Index の実行コスト
料金:キャッシュヒット・入力・出力
コンテキストウィンドウ
コンテキストウィンドウ
速度
出力速度(1秒あたりのトークン数)で測定
出力速度
Intelligence Indexのタスクあたりの時間
遅延
最初のトークンまでの時間(秒)で測定
遅延: 最初の回答トークンまでの時間
エンドツーエンド応答時間
Seconds to output 500 tokens, calculated based on time to first token, 'thinking' time for reasoning models, and output speed
エンドツーエンド応答時間
よくある質問
Claude Fable 5.1 (Max, Default Fallback)の知能が上です。Artificial Analysis Intelligence Indexで、Claude Fable 5.1 (Max, Default Fallback)のスコアは53、Qwen3.8 Max (0902)は45です。
Claude Fable 5.1 (Max, Default Fallback)の方が高速です。出力速度はClaude Fable 5.1 (Max, Default Fallback)が毎秒62.5トークン、Qwen3.8 Max (0902)が毎秒37.6トークンです。
Qwen3.8 Max (0902)の方が安価です。100万トークンあたりの料金はQwen3.8 Max (0902)が$1.18、Claude Fable 5.1 (Max, Default Fallback)が$7.17です(キャッシュヒット/入力/出力を7:2:1とした場合)。
Qwen3.8 Max (0902)の方が遅延が短くなっています。最初のトークンまでの時間はQwen3.8 Max (0902)が2.69秒、Claude Fable 5.1 (Max, Default Fallback)が259.51秒です。
Claude Fable 5.1 (Max, Default Fallback)のコンテキストウィンドウが大きくなっています。Claude Fable 5.1 (Max, Default Fallback)は1.0Mトークン、Qwen3.8 Max (0902)は980kトークンに対応しています。