MicroEvals
Public evaluations
Showing 2721-2740 of 2992



Ok

hey chat r we cooked











001



A single-turn comparison of Claude Opus 4.8, GPT-5.6 Luna, GPT-5.6 Sol, Gemini 3.1 Pro Preview, and Gemini 3.6 Flash across practical assistant capabilities: tool honesty, quantitative reasoning, scheduling, Bayesian inference, evidence synthesis, coding, infrastructure design, constrained writing, ambiguity handling, optimization, policy interpretation, and multilingual grounding. Prompt 1 is a web-access diagnostic and should be scored separately. Prompts 2–12 are self-contained and should not require browsing. The comparison evaluates each model at the reasoning setting shown by Artificial Analysis, so results represent best-configured endpoint quality rather than equalized compute, latency, or cost.

