All articles

September 7, 2026

OpenBMB's MiniCPM5-2B scores 15 on the Artificial Analysis Intelligence Index v4.2, the highest of any open weights model under 4B total parameters

See model page

OpenBMB is the open-source AI group behind the MiniCPM series of efficient small models. MiniCPM5-2B is a 2.6B parameter dense reasoning model with text input and output, released under Apache 2.0.

Scoring 15 on the Intelligence Index, MiniCPM5-2B sits one point behind Ling 3.0 Tiny (16), which has ~3x the total parameters. Among open weights models under 4B total parameters, the next best score is Granite 4.2 3B (11).

Key results:

The highest Intelligence Index of any open weights model under 4B total parameters, setting a new Pareto-optimal point on Intelligence vs. Total Parameters: Its score of 15 is 4 points clear of Granite 4.2 3B (11). With 2.6B total parameters, it is 1 point ahead of Qwen3.5 4B (Reasoning, 14, estimated) with 44% fewer parameters, and level with Qwen3.5 9B (Reasoning, 15, estimated) at roughly 4x its size. As a dense model, its size advantage is in memory footprint rather than active-parameter compute.

Strong agentic performance at this size: Its GDPval-AA v2 Elo of 831 leads <4B models, and on τ³-Banking it is joint-first with Ling 3.0 Tiny at 21%, compared to 8% for the next best model, Granite 4.2 8B. On AA-Briefcase, it placed second among the measured models in the comparison set with an Elo of 438, above Granite 4.2 8B (324) and just below Ling 3.0 Tiny (485).

Knowledge, coding and long context are where it gives ground. MiniCPM5-2B places 7th in the set on Humanity's Last Exam (9%, behind Gemma 4 12B (Reasoning) at 16%), 8th on Terminal-Bench v2.1 (9%, behind Qwen3.5 9B (Reasoning) at 29%) and scores 0% on CritPt. On SciCode it is second of the five measured models at 26%, behind Granite 4.2 8B (31%). On AA-LCR v1.1 it scores 59%, 5th in the set, one point behind Ling 3.0 Tiny (60%). On GDP.pdf, our new professional document reasoning evaluation, it passes 1% of tasks outright, behind gpt-oss-20b (high) at 2%.

Its AA-Omniscience score of -12 is earned by abstaining from answering rather than accuracy. MiniCPM5-2B attempts only 29% of AA-Omniscience questions, giving it a Non-Hallucination Rate of 78%. Its accuracy of 8% is a point below Ling 3.0 Tiny (9%) and half that of Qwen3.5 9B (Reasoning, 16%). Peers that attempt far more questions are penalized heavily, with Qwen3.5 9B (Reasoning) at -53 and gpt-oss-20b (high) at -63.

It is token-efficient for a reasoning model. MiniCPM5-2B used 19k output tokens per Intelligence Index task, joint-lowest in the comparison model set with Granite 4.2 3B (19k). Ling 3.0 Tiny spends 56k, roughly 3x as many, for 1 more index point.

Additional model details:

Parameters: 2.6B (dense)

Context window: 131k tokens

Input modalities: Text only

License: Apache 2.0

Agentic capability is MiniCPM5-2B's edge at this size. On AA-Briefcase, our agentic knowledge work evaluation, its Elo of 438 is second in the set behind Ling 3.0 Tiny (485) and ahead of Granite 4.2 8B (324). The Agentic Index is the weighted average of the three agentic evaluations in the Intelligence Index: AA-Briefcase, GDPval-AA v2 and τ³-Banking.

On GDPval-AA v2, which tests models on real-world work tasks against a human baseline of 1,000, MiniCPM5-2B reaches an Elo of 831, ~110 points ahead of Ling 3.0 Tiny (718) and ~180 ahead of Granite 4.2 8B (647). Models at this scale usually sit far lower, with LFM2.5-2.6B at 204 and Gemma 4 E4B (Reasoning) at 178.

MiniCPM5-2B is also token-efficient for a reasoning model. It uses 19k output tokens per Artificial Analysis Intelligence Index task, 11k of them reasoning tokens, against 56k for Ling 3.0 Tiny and 33k for Granite 4.2 8B. Output token use matters for the on-device and edge deployments a 2.6B model targets.

Full results across the Artificial Analysis Intelligence Index evaluations. Arrows compare MiniCPM5-1B (Reasoning) with MiniCPM5-2B where both have results; otherwise they highlight MiniCPM5-2B.

MiniCPM5-2B scored 23 on the recently updated Artificial Analysis Intelligence Index v4.1.1. Its new score of 15 on v4.2 reflects the updated evaluation mix and weightings, and scores across the two versions are not directly comparable.

The weights for MiniCPM5-2B are available under an Apache 2.0 licence at https://huggingface.co/openbmb/MiniCPM5-2B