MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Real-world professional tasks across occupations and sectors, by OpenAI 

Loading questions…

Public evaluations

Showing 401-420 of 10591

Best Website
👍1
Best Website

Which makes the best website

You work as a Senior Lifestyle Manager at a luxury concierge...
👍1
You work as a Senior Lifestyle Manager at a luxury concierge...

Line dot circle
👍1
Line dot circle

Crée un mini-jeu 3D jouable dans le navigateur en un seul fi...
👍1
Crée un mini-jeu 3D jouable dans le navigateur en un seul fi...

AI Model Comparison Test
👍1
AI Model Comparison Test

Compare multiple AI models on accuracy, reasoning, coding, mathematics, and general problem-solving ability using the same prompts.

Medical Quiz 3
👍1
Medical Quiz 3

Cardiac phisiology quiz.

Extensively give me an Islamic Table of Contents in 80 parts...
👍1
Extensively give me an Islamic Table of Contents in 80 parts...

CloudOps Architect Benchmark
👍1
CloudOps Architect Benchmark

A five-scenario benchmark designed to evaluate an LLM’s ability to perform as a senior Cloud Architect, DevOps Engineer, and SRE. The test covers production Kubernetes architecture, incident response, cloud security, legacy-to-cloud migration, and FinOps. It focuses on technical correctness, practical decision-making, trade-off reasoning, scalability, reliability, and depth of Cloud and Kubernetes expertise.

👍1
Just a test

Hshsjshs

👍1
Building an OS from scratch

Evaluate AI models on their ability to design, implement, debug, and explain an operating system from scratch. The evaluation covers boot processes, kernel architecture, C and assembly programming, memory management, interrupts, scheduling, system calls, device drivers, and debugging low-level code. Models should demonstrate technical correctness, sound engineering judgment, coherent implementation plans, and the ability to produce buildable, testable code. Reward explicit assumptions, incremental development, reproducible testing, security-conscious design, and honest acknowledgment of limitations. Penalize hallucinated APIs, incompatible code, unverified claims, and unnecessarily complex implementations. The goal is to determine which model can act as a reliable systems-programming mentor and engineering partner throughout the development of a small operating system.

Guess a number
👍1
Guess a number

# Prompt de investigación: Nueva arquitectura/algoritmo para...
👍1
# Prompt de investigación: Nueva arquitectura/algoritmo para...

uyftu
👍1
uyftu

API Key for "diqddi88@gmail.com"

1111
👍1
1111

Grok 4.3 high
👍1
Grok 4.3 high

Double Pendulum Simulation
👍1
Double Pendulum Simulation

Integral
👍1
Integral

Test 1
👍1
Test 1

Test 1

Two Math challenges
👍1
Two Math challenges

Persian prompts but translated in output

👍1
yooooo