MicroEvals
Public evaluations
Showing 81-100 of 6437
The purpose of this is to evaluate how good various AI models are at a variety of Minecraft skills like planning, designing, puzzles, and providing accurate information. It’ll also test to see how good each AI is at coding a web app clone of the game.



My goal is to determinate what level qwen3.8:27b actually is, and is it worth running locally by making them code the same html browser game.

Physics + Math + Programming + Scientific Reasoning

Coding and visual perfomance

Tests whether models can distinguish clean version-control state from actual data safety, identify unsupported assumptions, classify recoverability correctly, and propose practical safeguards without overstating certainty.


sfvsfvsvsvsv







Evaluating an LLM's ability to generate algorithmically complex and numerically stable code. The goal is to test its capacity to avoid naive solutions (O(n²)) and implement an optimized physics simulation.


Multitask development

5 challenging math problem