MicroEvals
Public evaluations
Showing 261-280 of 6466
How well can AI create fun, little games as web apps that can be played on mobile devices? This benchmark tests AI’s abilities to generate good mobile UI and controls as well as basic gameplay experiences.


local test



Which makes the best website




Compare multiple AI models on accuracy, reasoning, coding, mathematics, and general problem-solving ability using the same prompts.

Cardiac phisiology quiz.

A five-scenario benchmark designed to evaluate an LLM’s ability to perform as a senior Cloud Architect, DevOps Engineer, and SRE. The test covers production Kubernetes architecture, incident response, cloud security, legacy-to-cloud migration, and FinOps. It focuses on technical correctness, practical decision-making, trade-off reasoning, scalability, reliability, and depth of Cloud and Kubernetes expertise.



API Key for "diqddi88@gmail.com"




Test 1