MicroEvals
Public evaluations
Showing 61-80 of 2486




The purpose of this is to evaluate how good various AI models are at a variety of Minecraft skills like planning, designing, puzzles, and providing accurate information. It’ll also test to see how good each AI is at coding a web app clone of the game.

Physics + Math + Programming + Scientific Reasoning

Tests whether models can distinguish clean version-control state from actual data safety, identify unsupported assumptions, classify recoverability correctly, and propose practical safeguards without overstating certainty.


sfvsfvsvsvsv

Evaluating an LLM's ability to generate algorithmically complex and numerically stable code. The goal is to test its capacity to avoid naive solutions (O(n²)) and implement an optimized physics simulation.

Multitask development

5 challenging math problem





Generate an SVG of a pelican riding a bicycle
