comparison of model for use case
comparison of model for use case
Prompt
see: I'll pull all four comparison pages first.The first page returned a default model set rather than the four models in the URL. I'll check whether the other three do the same.Same default set again. I'll check the last two, then look for the requested models' data another way.Three of four are identical. I'll confirm the fourth and search for the requested models' actual numbers in parallel.That OfficeChai piece has real Sonnet 5.5 data. I'll look for GPT-6.1 Sol and the effort-level variants next.GPT-6.1 Sol numbers are in. Now the Claude 5.5 xhigh variants and GPT-6 Astra xhigh.Sonnet 5.5 at xhigh scores 52 versus 56 at max, a notable gap. I'll pull the full model pages and the remaining variants to get cost and token detail.The model pages expose summary stats, which is what I need. I'll fetch the remaining pages and find the URLs for Opus 5.5 and GPT-6 Astra variants.Sonnet 5.5 Max: 56 at $7.67/task versus xhigh's 52 at $2.75. Now the GPT-6.1 Sol pages for pricing and speed, plus searches for the Opus 5.5 and Astra variants.Sol Max: 52 at $0.72/task but slow (55.5 t/s). Now the Sol xhigh page and searches to surface the Opus 5.5 and Astra pages.Sol xhigh: 51 at $0.39/task. Two variants still missing: Opus 5.5 xhigh and GPT-6 Astra xhigh.Opus 5.5 scales 42 → 51 → 54 → 56 → 58 from low to max. I'll fetch the Opus xhigh page and find the GPT-6 Astra pages.Opus xhigh matches Sonnet Max's 56 at less than half the cost. I'll confirm Opus High's cost directly and find the GPT-6 Astra pages.Search results here mix snapshots from older Index versions (Astra Max shows as 61, then 55, then 53), so those aren't safe to use. I'll fetch the live Astra xhigh page and Opus High directly.I now have all eight variants. One more fetch: Opus 5.5 High's actual cost, since it bears on the Opus-versus-Sonnet question.On the Artificial Analysis (AA) Intelligence Index, these variants are much closer on capability than on cost. Opus 5.5 Max (58) is the only clear leader, and everything else lands between 51 and 56. Cost per task is where they separate. GPT-6.1 Sol is 3-4x cheaper than GPT-6 Astra or Sonnet 5.5 at xhigh, at the same score. Sonnet 5.5 Max is the weakest value in the set, because Opus 5.5 xhigh matches its 56 for under half the cost. Two notes on the data: - **Links:** All four of your links returned AA's default view when I fetched them (Opus 5.5 Max, GPT-6 Astra Max, Gemini 4 Argon High), probably because the model selection is applied in the browser. So I pulled each variant's own AA model page instead. All figures are AA's own, Index v4.3.2, and I haven't verified them independently. Unsuffixed slugs are the Max setting. - **Disclosure:** I'm Claude Sonnet 5.5, so I'm one of the models being compared. I've followed the numbers where they go, and at Max they aren't flattering to Sonnet. | Variant | Index | $/task | Output tokens (full index) | Tokens/s | TTFT | |---|---|---|---|---|---| | Opus 5.5 Max | 58 | $5.98 | 260M | 96 | 724s | | Opus 5.5 xhigh | 56 | $3.46 | 100M | 77 | 136s | | Sonnet 5.5 Max | 56 | $7.67 | 420M | 125 | 442s | | Opus 5.5 High* | 54 | $1.82 | 53M | 73 | 33s | | GPT-6 Astra Max | 53 | $3.26 | 60M | 45 | 377s | | GPT-6.1 Sol Max | 52 | $0.72 | 67M | 56 | 327s | | GPT-6 Astra xhigh | 52 | $2.31 | 38M | 46 | 216s | | Sonnet 5.5 xhigh | 52 | $2.75 | 110M | 111 | 36s | | GPT-6.1 Sol xhigh | 51 | $0.39 | 36M | 55 | 134s | *Not in your links; added as the useful Opus reference point. TTFT (time to first token) includes thinking time. List prices per 1M tokens (in/out): Opus $4/$20, Sonnet 5.5 and Sol $2/$10, Astra $10/$50. **Below the top, capability is close to a tie.** I'd treat 51-54 as one band. I saw no confidence intervals on the Index in the page text, and two of its ten components (SciCode, CritPt) are flagged "Under review." **Token volume, not list price, drives cost.** Sol and Sonnet 5.5 have identical $2/$10 pricing, yet Sol Max costs $0.72 per task against $2.75 for Sonnet xhigh at the same score of 52. Sol emits fewer tokens (67M vs 110M). AA's cache discounts (95% vs 90%) also imply cached input of $0.10/M for Sol versus $0.20/M for Sonnet, which is the same as Opus's. I can't see AA's cost breakdown in the page text, so that cache point is inference. Sonnet Max's 420M tokens take it to $7.67. **On this index, Opus is the cheaper way to buy a score within Anthropic's lineup.** Opus xhigh matches Sonnet Max's 56 at about 45% of the cost and under a quarter of the tokens. Opus High (54, $1.82) beats Sonnet xhigh (52, $2.75) on score, cost and tokens, with a similar ~35s wait to first token. Sonnet's genuine edge is decode speed (111-125 tokens/s vs 73-96 for Opus). But speed per token isn't speed per task: AA's time-per-task has Opus Max slower than Astra Max (783s vs 601s) despite decoding twice as fast, because it writes about four times the tokens. **Astra's premium is hard to justify here.** Astra Max is one point above Sol Max at 4.5x the cost. Astra xhigh matches Sol Max's 52 at 3.2x the cost, and its one edge is a shorter wait to first token (216s vs 327s). Both OpenAI families are also the slowest per token (45-56 tokens/s). **Max settings are batch settings, and they pay off unevenly.** Going from xhigh to Max adds 1 point on Sol and Astra, 2 on Opus and 4 on Sonnet, at 1.4-2.8x the cost. It also stretches time to first token to roughly 5-12 minutes. Even xhigh sits at 2-4 minutes for Opus, Sol and Astra. Only Sonnet xhigh and Opus High are near 35s, against a reasoning-model median around 4s, so none of the top settings suits interactive use. **Where Opus Max's lead comes from.** The one pair I could split by benchmark is Opus Max vs Astra Max, from AA's default table. - **Opus leads** on the two agentic knowledge-work evals: GDPval-AA Elo 1866 vs 1542 and AA-Briefcase 1807 vs 1569. After AA's (Elo-500)/2000 scaling, that's about 68 vs 52 and 65 vs 53. - **It also leads** on SciCode (67 vs 56) and Humanity's Last Exam (61 vs 55). - **Terminal-Bench 4.0** is level (60 vs 59), and Astra wins GDP.pdf (31 vs 26). So the five-point gap is mostly document-style work and scientific coding, while terminal-based agentic coding is a wash. AA's per-benchmark charts didn't come through as text for the other variants. If you paste that panel, I can extend this. **Caveats.** This is early data: Sonnet 5.5 and Sol launched about ten days ago. AA tested Sonnet 5.5 on an early build that Anthropic later found mishandled requests using structured outputs, and a third-party write-up says AA plans to re-run the affected evaluations on the public release, so Sonnet's numbers may move. AA also re-versions the Index often: older cached AA pages list Astra Max at 61, then 55, then 53 as the Index version changed. Don't mix scores across versions or snapshots. And AA's tasks are agentic-heavy, so cost per task reflects long tool-using loops, and chat-style usage will look different. On this data, Sol is the standout for capability per dollar on agentic work, and Opus at High or xhigh is the sensible ceiling before Max. Sonnet 5.5 earns a place mainly at xhigh, when decode speed matters. Astra needs a reason beyond this index to justify its price. Question: im building an agentic workflow. im limited to gpt-6 sol/high as my archictect (GPT chat best offering), who creates the plan and implementation task packets - these can be reviewed and enhanced by any model/reasoning, any model/reasoning can implement, any model/reasoning can QA, any model/reasoning can serve as the escalation point. the agentic workflow will first build me a trading bot and do strategy discovery, again all model/reasoning options are open bar architect. i have 2 GPT plus accounts and one Claude pro account. suggest the best configuration. does it change per task?