Strategy & Ops Index
Assesses model performance across the strategy and operations domain. Capabilities evaluated include domain-specific knowledge (business and management, accounting, corporate and markets), strategy and planning, customer support, records management, and more.
查看代表性工作流The Artificial Analysis Strategy & Ops Index combines performance across benchmarks chosen for strategy, operations, and office administration. We map common tasks from O*NET occupational classifications, then select benchmarks that represent this real-world work. Weights are derived from how often capabilities appear across those tasks.
This composite metric provides a single score for tracking model performance across operations and administrative work. 所有底层基准测试均由 Artificial Analysis 独立运行。有关评测的实施方式,请参阅我们的智能基准测试方法论。
| 能力 | 权重 | 评测 |
|---|---|---|
| Agentic Knowledge Work | 35% | GDPval-AA v2和AA-Briefcase |
| Business Knowledge | 30% | AA-Omniscience Business Accuracy |
| Agentic Tool Use | 30% | AutomationBench-AA Operations, HR, Marketing, Sales & Support |
| Long-Context | 5% | LCR和GDP.pdf |
得分
Artificial Analysis Strategy & Ops Index
Artificial Analysis Strategy & Ops Index:能力明细
能力明细
Artificial Analysis Strategy & Ops Index:Business Knowledge
代表性工作流
这些真实工作流重点检验 Strategy & Ops Index 中权重最高的能力。
示例:Assess a mid-market SaaS company's competitive position from market-share data, analyst reports, and win/loss notes, work through a Porter's Five Forces and SWOT read, and formulate three prioritized strategic options with the trade-offs of each.
示例:Scan 3,000 invoices of different formats before month-end and extract line-items into structured fields to add to the general ledger.
示例:Absorb a 40% support surge after a product recall in a CRM ticketing queue with 20-minute hold times to triage incoming tickets by SLA, de-escalate frustrated customers in live chat, and follow the approved recall script verbatim.
示例:Reconcile an executive's schedule when they're triple-booked across a full week of calendars in a scheduling tool, including external stakeholders with limited availability, to weigh free/busy windows, propose conflict resolutions ranked by stakeholder seniority, and draft rescheduling notes.
示例:Clean up a stalled month-end close where multiple departments coded the same expenses to different GL accounts across hundreds of invoices to propose a consistent chart-of-accounts mapping, identify entries needing reclassification, and draft adjusting journal entries.
示例:Consolidate five years of training records split across paper files and two unindexed document systems for a regulatory request to build one audit-ready manifest, flag missing records with supporting evidence, and propose a retention schedule for the next cycle.
成本
Artificial Analysis Strategy & Ops Index:单任务成本
Artificial Analysis Strategy & Ops Index 与单任务成本
速度
Artificial Analysis Strategy & Ops Index:单任务耗时
输出 token
Artificial Analysis Strategy & Ops Index:单任务输出 token
发布日期
Artificial Analysis Strategy & Ops Index 与发布日期
常见问题
根据 Artificial Analysis Strategy & Ops Index,目前在战略与运营工作上表现最佳的 AI 模型是 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) (60)、Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) (58)和GPT-6 Astra (max) (58)。新模型发布后,排行榜会随之更新。
有。Artificial Analysis Strategy & Ops Index 是一项独立基准测试,用于衡量 AI 模型在战略与运营工作上的表现。它评估商业知识、智能体知识工作、在业务应用中的智能体工具使用以及长上下文分析等能力。
Strategy & Ops Index 是 Artificial Analysis 推出的综合基准测试,用于评估模型在战略与运营领域的表现。评估能力包括商业与管理、会计、企业与市场等专业知识,以及战略与规划、客户支持、档案管理等。
Strategy & Ops Index 按各项能力子分数的加权平均值计算。各项子分数及其权重为:Business Knowledge (30%)、Agentic Knowledge Work (35%)、Agentic Tool Use (30%)和Long-Context (5%)。
Strategy & Ops Index 包含 AA-Omniscience Business Accuracy、GDPval-AA v2、AA-Briefcase、AutomationBench-AA Operations, HR, Marketing, Sales & Support、LCR和GDP.pdf。
在已公布结果的模型中,Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) 目前以 60 分位居 Strategy & Ops Index 榜首。 查看模型
Strategy & Ops Index 得分越高,表示模型在构成该指数的各项基准测试中整体表现越强。对于特定用例,单项基准测试结果可能比综合得分更具参考价值。