July 30, 2026
Inkling Small lands within a point of Inkling on the Artificial Analysis Intelligence Index with less than a third of the parameters
Thinking Machines' new Inkling Small scores 40 on the Artificial Analysis Intelligence Index, within a point of its flagship sibling Inkling with less than a third of the total and active parameters
Inkling Small is Thinking Machines Lab's second model release, arriving two weeks after Inkling launched at 41 on the Artificial Analysis Intelligence Index. Thinking Machines is the San Francisco-based AI lab founded by former OpenAI CTO Mira Murati. Inkling Small is an open weights reasoning model with 276B total parameters (12B active MoE), text, image, and speech input, and a 256K token context window.
Key results:
➤ Inkling Small holds a similar tier of intelligence to Inkling’s at less than one third of its size: 276B total parameters (12B active) vs. 975B (41B active) for Inkling. No open weights model at its size or smaller scores higher on the Intelligence Index. DeepSeek V4 Flash (max), at a similar size of 284B total (13B active) also scores 40, MiniMax-M3 reaches 44 with 23B active, while GLM-5.2 (max) reaches 51 with 40B active.
➤ Inkling Small meets or exceeds Inkling on several coding and frontier reasoning evaluations. It scores higher on Humanity's Last Exam (32% vs. 30%), GPQA Diamond (89% vs. 87%), CritPt (8% vs. 5%), and SciCode (49% vs. 46%), and achieves the same score on Terminal Bench v2.1 (55%).
➤ Inkling Small is not as strong as Inkling on Agentic tasks and factual knowledge. Inkling Small trails Inkling on τ³-Banking (15% vs. 24%), though it edges ahead on GDPval-AA v2 (1269 vs. 1237 Elo). On the AA-Omniscience Index it scores -9 vs. the flagship's positive 2, meaning incorrect answers outweigh correct ones; this is driven by lower AA-Omniscience Accuracy (31% vs. 40%) rather than hallucination: Inkling Small’s Hallucination Rate is slightly lower than its larger sibling (57% vs. 63%).
➤ Inkling Small averaged ~24K output tokens per Intelligence Index task, slightly fewer than Inkling (~25K), while peers at its intelligence level averaged far more. DeepSeek V4 Flash averaged ~45K and GPT-5.4 mini (xhigh) ~78K. Running the full Intelligence Index took roughly the same number of output tokens for Inkling Small and Inkling.
Additional model details:
➤ Type: Open weights reasoning model (Apache 2.0 license)
➤ Size: 276B total parameters, 12B active (MoE)
➤ Input modalities: Text, image, and speech
➤ Output modalities: Text
➤ Context window: 256K tokens (Inkling supports 1M)
Congratulations to the team at Thinking Machines on the release!

Inkling Small is a 276B parameter MoE with 12B active parameters, less than a third of Inkling. No open weights model at its size or smaller scores higher. DeepSeek V4 Flash (max) has 13B active parameters (284B total) and achieves the same score, while MiniMax-M2.7 (230B, 10B active) sits 2 points lower.

Inkling Small scores lower than Inkling on AA-Omniscience, coming in at -9 where Inkling scores 2, meaning it has more incorrect answers than correct ones. The lower score is driven by lower accuracy (31% vs. 40%), which is typical of smaller total parameter count models. Inkling Small's Hallucination Rate is slightly lower (57% vs. 63%).

Inkling Small averaged ~24K output tokens per Artificial Analysis Intelligence Index task, slightly fewer than Inkling (~25K) - token efficient next to peers near its intelligence level like DeepSeek V4 Flash (max, ~45K) and GPT-5.4 mini (xhigh, ~78K). Running the full Intelligence Index evaluations took ~131M output tokens, slightly more than the flagship's ~128M.

Inkling Small reaches a higher score than its larger sibling on AA-Briefcase, our long-horizon agentic knowledge-work benchmark: 917 Elo vs. 839 for Inkling, driven significantly by better presentation. Rubric pass rates are nearly identical (20% vs. 19%), suggesting Inkling Small's outputs are better presented rather than substantially more correct; however, it also finished tasks in less than half the turns (34 vs. 81 per task on average).

Full breakdown of Inkling Small's performance across the nine evaluations in the Artificial Analysis Intelligence Index:

See Artificial Analysis for further details and benchmarks: https://artificialanalysis.ai/models/inkling-small
Read the latest

Agnes AI releases Agnes 2.5 Pro Alpha
Agnes 2.5 Pro Alpha
July 29, 2026

Claude Opus 5: the new leader in agentic knowledge work
Claude Opus 5 is the new leader on our agentic knowledge work benchmark, AA-Briefcase, outperforming Claude Fable 5 by nearly 150 Elo while reducing Cost per Task by 20%
July 24, 2026

Opus 5: Fable 5 level intelligence at a lower cost per task
Claude Opus 5 is narrowly the most intelligent model on the Artificial Analysis Intelligence Index, offering comparable intelligence to Fable 5 at 26% lower Cost per Task
July 24, 2026