AIの独立分析
AIの全体像を把握し、ユースケースに最適なモデルとプロバイダーを選択
注目情報
知能
独自評価に基づく主要AIモデルの知能
Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index by Open Weights / Proprietary
Cost per Intelligence Index Task
Intelligence Index vs. Cost per Intelligence Index Task
最先端言語モデルの知能の推移
主要なコーディングエージェントの、エンドツーエンドのソフトウェアエンジニアリングタスクにおける性能、コスト、実行時間
Artificial Analysis Coding Agent Index
画像と動画
Image ArenaとVideo Arenaのリーダーボード上位モデル(95%信頼区間付き)
テキストから画像生成のリーダーボード
音声
Text to Speech Arena、音声文字起こし、音声対話の各評価における上位モデル
Text to Speech Arena Leaderboard
Artificial Analysis Agentic Index
Intelligence Evaluations
Agentic real-world work tasks, (Elo-500)/2000
Agentic tool use
Agentic coding & terminal use
Coding
Reasoning & knowledge
Scientific reasoning
Physics reasoning
Knowledge
1 - hallucination rate
Long context reasoning
Agentic knowledge work, Elo
Agentic SaaS workflows
Legal agentic work, criterion pass rate
Agentic business operations
Instruction following
Long-horizon agentic tasks
Kubernetes incident root-cause analysis
Visual reasoning
AA-Briefcase
AA-Briefcaseは、長期的な知識労働を対象とした最先端のエージェント型評価です。スプレッドシート、プレゼンテーション、メモなどの成果物を必要とする現実的なビジネスワークフローでエージェントをテストします
AA-Briefcase Elo
AA-Omniscience
AA-Omniscienceは、正確な回答を評価し、誤った推測にはペナルティを与える、知識とハルシネーションのベンチマークです。さまざまな分野で事実に基づく信頼性の高い出力を生成するモデルを包括的に把握できます
AA-Omniscience Index
GDPval-AA v2
GDPval-AA v2は、幅広い職業における現実の経済的価値が高いタスクでAIモデルを評価します
GDPval-AA v2 Leaderboard
Artificial Analysis Openness Indexは、さまざまな構成要素の利用可能性と透明性に基づき、モデルがどの程度「オープン」であるかを評価します。
Artificial Analysis Openness Index: Components
Artificial Analysis Openness Index vs. Artificial Analysis Intelligence Index
出力トークン
独自評価に基づく主要AIモデルの出力トークン数
Output Tokens per Intelligence Index Task
コスト
独自評価に基づく主要AIモデルの料金と実環境でのコスト
Cost per Intelligence Index Task
Cost to Run Artificial Analysis Intelligence Index
Pricing: Cache Hit, Input, and Output
速度と遅延
ファーストパーティーAPIの性能比較