MicroEvals
Public evaluations
Showing 401-420 of 10591

Which makes the best website




Compare multiple AI models on accuracy, reasoning, coding, mathematics, and general problem-solving ability using the same prompts.

Cardiac phisiology quiz.


A five-scenario benchmark designed to evaluate an LLM’s ability to perform as a senior Cloud Architect, DevOps Engineer, and SRE. The test covers production Kubernetes architecture, incident response, cloud security, legacy-to-cloud migration, and FinOps. It focuses on technical correctness, practical decision-making, trade-off reasoning, scalability, reliability, and depth of Cloud and Kubernetes expertise.
Hshsjshs
Evaluate AI models on their ability to design, implement, debug, and explain an operating system from scratch. The evaluation covers boot processes, kernel architecture, C and assembly programming, memory management, interrupts, scheduling, system calls, device drivers, and debugging low-level code. Models should demonstrate technical correctness, sound engineering judgment, coherent implementation plans, and the ability to produce buildable, testable code. Reward explicit assumptions, incremental development, reproducible testing, security-conscious design, and honest acknowledgment of limitations. Penalize hallucinated APIs, incompatible code, unverified claims, and unnecessarily complex implementations. The goal is to determine which model can act as a reliable systems-programming mentor and engineering partner throughout the development of a small operating system.



API Key for "diqddi88@gmail.com"





Test 1

Persian prompts but translated in output