Strategy & Ops Index
Assesses model performance across the strategy and operations domain. Capabilities evaluated include domain-specific knowledge (business and management, accounting, corporate and markets), strategy and planning, customer support, records management, and more.
代表的なワークフローを見るThe Artificial Analysis Strategy & Ops Index combines performance across benchmarks chosen for strategy, operations, and office administration. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across operations and administrative work. 基礎となるすべてのベンチマークは、Artificial Analysisが独立して実施しています。評価方法の詳細は、知能ベンチマークの方法論をご覧ください。
| 能力 | 重み | 評価 |
|---|---|---|
| Agentic Knowledge Work | 35% | GDPval-AA v2、AA-Briefcase |
| Business Knowledge | 30% | AA-Omniscience Business Accuracy |
| Agentic Tool Use | 30% | AutomationBench-AA Operations, HR, Marketing, Sales & Support |
| Long-Context | 5% | LCR、GDP.pdf |
スコア
Artificial Analysis Strategy & Ops Index
Artificial Analysis Strategy & Ops Index:能力の内訳
能力の内訳
Artificial Analysis Strategy & Ops Index:Business Knowledge
代表的なワークフロー
Strategy & Ops Indexで特に重視される能力を測る、実際の業務を想定したワークフローです。
例:Assess a mid-market SaaS company's competitive position from market-share data, analyst reports, and win/loss notes, work through a Porter's Five Forces and SWOT read, and formulate three prioritized strategic options with the trade-offs of each.
例:Scan 3,000 invoices of different formats before month-end and extract line-items into structured fields to add to the general ledger.
例:Absorb a 40% support surge after a product recall in a CRM ticketing queue with 20-minute hold times to triage incoming tickets by SLA, de-escalate frustrated customers in live chat, and follow the approved recall script verbatim.
例:Reconcile an executive's schedule when they're triple-booked across a full week of calendars in a scheduling tool, including external stakeholders with limited availability, to weigh free/busy windows, propose conflict resolutions ranked by stakeholder seniority, and draft rescheduling notes.
例:Clean up a stalled month-end close where multiple departments coded the same expenses to different GL accounts across hundreds of invoices to propose a consistent chart-of-accounts mapping, identify entries needing reclassification, and draft adjusting journal entries.
例:Consolidate five years of training records split across paper files and two unindexed document systems for a regulatory request to build one audit-ready manifest, flag missing records with supporting evidence, and propose a retention schedule for the next cycle.
コスト
Artificial Analysis Strategy & Ops Index:タスク当たりのコスト
Artificial Analysis Strategy & Ops Index vs. タスク当たりのコスト
速度
Artificial Analysis Strategy & Ops Index:タスク当たりの時間
出力トークン
Artificial Analysis Strategy & Ops Index:タスク当たりの出力トークン
リリース日
Artificial Analysis Strategy & Ops Index vs. リリース日
よくある質問
Artificial AnalysisのStrategy & Ops Indexによると、現在戦略・オペレーション業務で最も高い性能を示すAIモデルはClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (60)、Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) (58)、GPT-6 Astra (max) (58)です。ランキングは新しいモデルのリリースに合わせて更新されます。
はい。Artificial AnalysisのStrategy & Ops Indexは、戦略・オペレーション業務におけるAIモデルの性能を測る独立したベンチマークです。ビジネス知識、エージェント型ナレッジワーク、業務アプリでのエージェント型ツール使用、長文コンテキスト分析を評価します。
Strategy & Ops Indexは、戦略・オペレーション分野におけるモデルの性能を評価するArtificial Analysisの複合ベンチマークです。ビジネスと経営、会計、企業、市場などの専門知識に加え、戦略と計画、顧客サポート、記録管理などを評価します。
Strategy & Ops Indexは、各能力のサブスコアを加重平均して算出します。サブスコアとその重みは、Business Knowledge (30%)、Agentic Knowledge Work (35%)、Agentic Tool Use (30%)、Long-Context (5%)です。
Strategy & Ops IndexにはAA-Omniscience Business Accuracy、GDPval-AA v2、AA-Briefcase、AutomationBench-AA Operations, HR, Marketing, Sales & Support、LCR、GDP.pdfが含まれます。
結果が公開されているモデルの中では、現在Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)がStrategy & Ops Indexで最高の60を記録しています。 モデルを見る
Strategy & Ops Indexのスコアが高いほど、指数を構成するベンチマーク全体での性能が優れていることを示します。特定の用途では、複合スコアよりも個別のベンチマーク結果のほうが参考になる場合があります。