August 27, 2026
Agnes AI's Agnes 2.5 Pro Beta scores 49 on the Artificial Analysis Intelligence Index, up 9 points from Agnes 2.5 Pro Alpha, driven by large agentic gains but using ~2x the output tokens
Agnes AI is a Singapore-based AI lab that trains full-modality foundation models and offers them through a free omni-modal API. Agnes 2.5 Pro Beta is a beta checkpoint of its next 2.5 Pro release, a text, image, and video input reasoning model with text output.
At 49 on the Intelligence Index, Agnes 2.5 Pro Beta moves Agnes from mid-pack to the frontier-adjacent tier, just behind Gemini 3.5 Flash (high, 52) and GPT-5.6 Luna (max, 52) and ahead of MiniMax-M3 (45).
Key results:
➤ Agnes 2.5 Pro Beta scores 49 on the Artificial Analysis Intelligence Index, a 9-point jump from Agnes 2.5 Pro Alpha (40). This places it just behind Gemini 3.5 Flash (high, 52) and GPT-5.6 Luna (max, 52), and ahead of MiniMax-M3 (45).
➤ Agentic capabilities drive the jump. The Artificial Analysis Agentic Index rises from 25 to 44, just behind Gemini 3.7 Flash (high, 45) and ahead of Gemini 3.5 Flash (high, 40) and MiniMax-M3 (36). τ³-Banking nearly triples from 12% to 36%, and GDPval-AA v2 rises from an Elo of 1171 to 1456 against a human baseline of 1000.
➤ Frontier reasoning evaluations improves more modestly. Humanity's Last Exam rises from 34% to 38%, GPQA Diamond from 88% to 91%, and CritPt from 11% to 16%.
➤ The AA-Omniscience improvement from -25 to -11 comes from abstention, not increased accuracy. Agnes 2.5 Pro Beta attempts only 45% of questions against Agnes 2.5 Pro Alpha's 94%, cutting the hallucination rate from 88% to 33%, but AA-Omniscience Accuracy halves from 33% to 17%.
➤ The intelligence gain required ~2x as many tokens from its predecessor. Agnes 2.5 Pro Beta uses 50k output tokens per Intelligence Index task, more than double Agnes 2.5 Pro Alpha's 24k.
Additional model details:
➤ Context window: 1M tokens
➤ Max output tokens: 65k
➤ Input modalities: Text and image
➤ Pricing: $0.10 / $0.30 / $0.01 per 1M input / output / cache hit tokens
➤ Availability: Agnes AI first-party API

Agnes 2.5 Pro Beta's largest gains are on agentic work. It achieves an Artificial Analysis Agentic Index score of 44, up from 25 from its predecessor. It scores an Elo of 1456 on GDPval-AA v2 against a human baseline of 1000, up from 1171 for Agnes 2.5 Pro Alpha, ahead of MiniMax-M3 (1384) and just behind GPT-5.5 (xhigh, 1489) and Gemini 3.7 Flash (high, 1527).

Agnes 2.5 Pro Beta's AA-Omniscience score improves from -25 to -11, however the improvement comes from abstention rather than accuracy improvement. It attempted only 45% of questions compared to 94% for Agnes 2.5 Pro Alpha, and while its hallucination rate improves from 88% to 33%, its AA-Omniscience Accuracy halves from 33% to 17%.

Agnes 2.5 Pro Beta uses 50k output tokens per Artificial Analysis Intelligence Index task, more than double Agnes 2.5 Pro Alpha's 24k. It uses more tokens than Qwen3.8 27B (xhigh, 47k) and GLM-5.3 (max, 41k).

Full results across the Artificial Analysis Intelligence Index:

Read the latest
Korean AI Lab Upstage releases Solar Mini 4
Korean AI Lab Upstage has released Solar Mini 4 which scores 24 on the Artificial Analysis Intelligence Index, but costs ~5x as much per task as GPT-6 Luna (max) despite similar per-token prices
September 30, 2026

Gemini 4 Argon: Google is back as one of the top three labs in intelligence achieved
Google’s new Gemini 4 Argon equals GPT-6 Astra on the Artificial Analysis Intelligence Index at 60% of the Cost per Task with discounted prices. Google is now back to being one of the top three labs in intelligence achieved
September 30, 2026

AA-AgentPerf-Local: Benchmarking local AI agents on laptops and workstations
Our open-source tool for testing how fast agentic AI runs on laptops and workstations, with launch results for the DGX Spark, Ryzen AI Halo, MacBook Pro (M5 Pro) and RTX 5090 across four open-weights models
September 29, 2026