All articles

August 10, 2026

Muse Glimmer: Benchmarks and Analysis

Meta returns to open weights: Muse Glimmer, its first open-weights release since Llama 4, scores 35 on the Artificial Analysis Intelligence Index. It is a 30B-parameter model, and the first from Meta to be released under Apache 2.0

Muse Glimmer (high) arrives 16 months after Llama 4, scoring 21 points above Llama 4 Maverick (14), Meta's last open weights release. It sits alongside Kimi K2.5 (Reasoning, 36) and just behind Qwen3.6 27B (Reasoning, 38) and Ling 3.0 Flash (38), and creates a two-tier Meta lineup together with the proprietary flagship Muse Spark 1.2 (xhigh, 57). Meta shared access with us ahead of public release for benchmarking.

Key takeaways:

➤ Meta's open-weights line is back, under its most permissive license yet: Every prior Meta open release shipped under a Llama License; Muse Glimmer uses Apache 2.0, placing almost no restrictions on commercial use or derivatives. We have evaluated the initial release at 44 on the Artificial Analysis Openness Index, our measure of model availability and transparency across model weights, training data, and other factors - equal to DeepSeek V4 Flash (0731), GLM-5.2, and Ling 3.0 Flash, and ahead of most open models. It also improves on Llama 4 Maverick's openness score, owing to its more permissive license and more extensive methodology disclosure.

➤ Strong intelligence for its parameter count: At 30B parameters, Muse Glimmer scores 5 points above Gemma 4 31B (Reasoning, 30) at the same size, and effectively matches 1T total parameter Kimi K2.5 (Reasoning, 36) with 33x fewer parameters. Qwen3.6 27B (Reasoning, 38) remains ahead on the Intelligence vs Parameters frontier at a slightly smaller size.

➤ Small enough to self-host on a single GPU, even at full context: Muse Glimmer is a 30B dense model (including a ~1.8B vision encoder, scoring 74% on the MMMU-Pro visual reasoning benchmark) with weights at ~60 GB in BF16 and ~18 GB in 4-bit. It features a hybrid-attention mechanism with three sliding-window layers for every global layer, which holds KV cache memory use to ~1.8 GB (minimum) at its pre-extension 128K context. This means the model can run at full context on a single H100 at BF16 precision, or on a higher-spec MacBook or RTX 5090 at 4-bit, with more breathing room if the vision encoder is not required.

➤ Agentic knowledge work is its weakness relative to its size class: Muse Glimmer scores 953 Elo on GDPval-AA v2, below the 1,000 human baseline and behind other models at its intelligence level, including the similarly sized Qwen3.6 27B (Reasoning, 1141) and Gemini 3.5 Flash-Lite (1141). GDPval-AA v2 is our leading metric for agentic performance, measuring models on realistic knowledge work tasks in an agentic loop via our reference harness Stirrup. Knowledge calibration follows the same pattern: its AA-Omniscience Index of -33 is low for its intelligence level, driven by an 82% hallucination rate (Qwen3.6 27B: 49%) rather than accuracy, where it matches its peers. Agentic tool use is the exception, with Muse Glimmer scoring 24% on Tau3-Banking, ahead of Gemini 3.5 Flash-Lite (18%) and Qwen3.6 27B (17%), among the best in its class.

➤ Model details: Muse Glimmer is a 30B dense model with a 128K token context window (plus extension), released under Apache 2.0. At the time of release, Meta is not serving the model on their API; pricing and serving speed depend on third-party providers, and the weights are available on Hugging Face.

We have evaluated the initial release of Muse Glimmer at 44 on the Artificial Analysis Openness Index, a measure of model availability and transparency across model weights, training data, and other factors. This places the model at an equal position to DeepSeek V4 Flash (0731 version), GLM-5.2 and Ling 3.0 Flash, ahead of most open models. It also improves on Llama 4 Maverick's openness score, owing to its more permissive license and more extensive methodology disclosure.

At 30B parameters, Muse Glimmer sits near the Intelligence vs Parameters frontier for open weights models: 5 points above Gemma 4 31B (Reasoning) at the same size, effectively matching Kimi K2.5 (Reasoning) at 33x fewer total parameters, and just behind Qwen3.6 27B (Reasoning), the parameter-efficiency leader in this class.

Muse Glimmer's gaps against its class concentrate in agentic evaluations: 953 Elo on GDPval-AA v2 against 1141 for Qwen3.6 27B (Reasoning), 1141 for Gemini 3.5 Flash-Lite, and 1004 for Kimi K2.5 (Reasoning), with Terminal-Bench v2.1 (52%) also behind Qwen3.6 27B (61%). The hallucination gap follows: 82% on AA-Omniscience against 49% for Qwen3.6 27B and 34% for Flash-Lite (lower is better). Its size-twin Gemma 4 31B (Reasoning) performs worse than Muse Glimmer on all of these measures. The exception is agentic tool use, with Muse Glimmer scoring strongly on Tau3-Banking (24%), ahead of Gemini 3.5 Flash-Lite (18%) and Qwen3.6 27B (17%).

Full breakdown of the individual evaluations in the Artificial Analysis Intelligence Index:

See Artificial Analysis for further details and benchmarks of Muse Glimmer: https://artificialanalysis.ai/models/muse-glimmer