July 29, 2026
Agnes AI has launched Agnes 2.5 Pro Alpha, a low-priced reasoning model reaching 39 on the Artificial Analysis Intelligence Index at $0.45/$0.90 per 1M tokens
See model pageAgnes AI is a Singapore-based AI lab that trains its own full-modality foundation models in-house across text, image, and video, and offers them through a free omni-modal API that has passed 3 million users. Agnes 2.5 Pro Alpha is a text, image, and video input reasoning model with text output.
Agnes 2.5 Pro Alpha scores 39 on the Artificial Analysis Intelligence Index, placing it mid-pack among reasoning models we have tested and below the current proprietary frontier. Its defining characteristic is price: at $0.45 per 1M input tokens and $0.90 per 1M output tokens, Agnes 2.5 Pro Alpha is among the lowest-priced proprietary models at its intelligence level.
Key results:
➤ Agnes 2.5 Pro Alpha debuts at 39 on the Artificial Analysis Intelligence Index, its first appearance on our leaderboard. This places it mid-pack among current reasoning models.
➤ Coding is a relative strength for Agnes 2.5 Pro Alpha, scoring 58.8 on the Coding Index, near the top for its intelligence tier. This is above similarly priced models including DeepSeek V4 Flash (56.2), GPT-5.4 mini (56.1) and Qwen3.7 Plus (55.9), and sits just behind Nex-N2-Pro (59.1).
➤ On academic reasoning, Agnes 2.5 Pro Alpha is competitive for its tier, scoring 32% on Humanity's Last Exam and 88% on GPQA Diamond. Its Humanity's Last Exam result leads similarly priced proprietary models including GPT-5.4 mini (xhigh, 27%) and GPT-5.4 nano (xhigh, 26%), and its GPQA Diamond score is in line with peers near its Intelligence Index.
➤ At $0.45/$0.90 per 1M input/output tokens, Agnes 2.5 Pro Alpha is the second cheapest model at its intelligence level. Blended at 7:2:1 (cache-input-output) it costs $0.18 per 1M tokens, behind only DeepSeek V4 Flash (Reasoning, Max Effort) at $0.06 among models scoring 38-40 on the Artificial Analysis Intelligence Index. Open weights alternatives including DeepSeek V4 Pro (Reasoning, Max Effort) and MiMo-V2.5-Pro match that $0.18 price at higher intelligence.
Additional model details:
➤ Context window: 1M tokens.
➤ Pricing: $0.45 / $0.90 / $0.0038 per 1M input / output / cache hit tokens.
➤ Input modalities: Text, image, and video.
➤ Availability: Agnes AI first-party API.

Agnes 2.5 Pro Alpha is strong in coding relative to its price tier: its Coding Index of 58.8 is above similarly priced models including DeepSeek V4 Flash (56.2), GPT-5.4 mini (56.1) and Qwen3.7 Plus (55.9).

On GDPval-AA v2, which rates performance on real-world work tasks against a human baseline of 1,000, Agnes 2.5 Pro Alpha scores 1168, above the baseline and level with GPT-5.4 mini (xhigh, 1169), ahead of Qwen3.7 Plus (943) but behind DeepSeek V4 Flash (max, 1189).

Agnes 2.5 Pro Alpha used 22k output tokens per Artificial Analysis Intelligence Index task, 16k of them reasoning tokens, 51% fewer than DeepSeek V4 Flash (max, 45k) and 72% fewer than GPT-5.4 mini (xhigh, 78k).

Agnes 2.5 Pro Alpha's per-evaluation profile across the ten evaluations in the Artificial Analysis Intelligence Index, showing its stronger results on GPQA Diamond and Humanity's Last Exam alongside the areas where it currently trails.

Read the latest

Announcing the Artificial Analysis Search Index: Same Agent, Different Search
The Artificial Analysis Search Index benchmarks how search API providers perform on quality, cost, and speed when used by an agent. We compare different Search API providers across a series of search-related benchmarks using the same agentic setup.
August 18, 2026

Announcing Optima: create a custom benchmark for your use case
Optima is a new platform for benchmarking models on your own workloads. Build a benchmark from your own files, agent traces or coding environment, run it across leading models in a single click, and compare quality alongside cost per task and time per task.
August 13, 2026

Gemini 3.7 Flash: On the Intelligence vs. Time per Task Pareto frontier
Google has released Gemini 3.7 Flash, improving 4 points over Gemini 3.6 Flash and reaching the Intelligence vs. Time per Task Pareto frontier
August 13, 2026