MicroEvals
Public evaluations
Showing 301-320 of 7822


Compare multiple AI models on accuracy, reasoning, coding, mathematics, and general problem-solving ability using the same prompts.

Cardiac phisiology quiz.


A five-scenario benchmark designed to evaluate an LLM’s ability to perform as a senior Cloud Architect, DevOps Engineer, and SRE. The test covers production Kubernetes architecture, incident response, cloud security, legacy-to-cloud migration, and FinOps. It focuses on technical correctness, practical decision-making, trade-off reasoning, scalability, reliability, and depth of Cloud and Kubernetes expertise.



API Key for "diqddi88@gmail.com"





Test 1

Persian prompts but translated in output


More 't' s and a 'T' is added to confuse AI

це кольба


Models are tasked with implementing a specific type of less commmon rate limiter which must be compatible with a third party library aiohttp