May 20, 2026
Cohere launches open weights model Command A+ that achieves 37 on the Artificial Analysis Intelligence Index
See model pageThe release of Command A+ places Cohere in line with Claude 4.5 Haiku on the Intelligence Index, and just above NVIDIA Nemotron 3 Super and Gemini 3.1 Flash-Lite.
Key Takeaways:
➤ Command A+ ranks first on AA-Omniscience Non-Hallucination at 86%, ~3 percentage points ahead of the next-best model. Its AA-Omniscience Accuracy is 9%, so the headline AA-Omniscience score lands at -4, demonstrating a similar archetype to Claude 4.5 Haiku, where the model knows its limits
➤ On Cohere’s API, Command A+ (~281 output tokens per second) is faster than several comparable open-weights and small to mid-sized proprietary models (e.g., GPT-5.4 nano, Claude 4.5 Haiku, and Grok 4.3), but still slower than Gemini 3.1 Flash-Lite Preview, which outputs 304 tokens per second
➤ Command A+ trails its peer set on scientific reasoning (HLE ~11%, GPQA Diamond ~76%) and on coding (Terminal-Bench Hard ~25%, SciCode ~38%), consistent with gaps on the hardest science and agentic coding benchmarks
➤ It supports visual reasoning and scores 63% on MMMU-Pro (between Claude 4.5 Haiku at 59% and GPT-5.4 nano (xhigh) at 65%)

In our pre-release testing, Command A+ performed strongly on speed for its intelligence, reaching 281 output tokens per second. This reflects higher intelligence and speed than models such as gpt-oss-120b, but sits behind the new Pareto frontier established by Gemini 3.5 Flash

Amongst comparable models, Command A+ is competitive on intelligence vs. output tokens to run the Artificial Analysis Intelligence Index

On AA-Omniscience, our knowledge and hallucination evaluation, Command A+ fits a positive archetype where it knows its limitations: although it has relative low performance on AA-Omniscience Accuracy, it also hallucinates the least on AA-Omniscience, resulting in a fairly strong relative headline AA-Omniscience score of -4

Breakdown of individual evaluations:

See Artificial Analysis for further details and benchmarks: https://artificialanalysis.ai/models/command-a-plus
Read the latest

Agnes AI releases Agnes 2.5 Pro Beta
Agnes 2.5 Pro Beta
August 27, 2026

Intelligence at pocket scale: Benchmarking small models and mobile phones
Independent intelligence benchmarking of small language models on a set of evaluations chosen for mobile device use, launched alongside mobile phone inference benchmarking with Liquid AI. We evaluate the same quantized builds used on mobile phones, and performance is measured on real devices.
August 24, 2026

Announcing the Speech Agent Arena: Compare Speech agents in real world conversations
Announcing our new Speech Agent Arena, evaluating Speech to Speech models on real-world scenarios to analyze conversational preference and task success rate
August 24, 2026