MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Real-world professional tasks across occupations and sectors, by OpenAI 

Loading questions…

Public evaluations

Showing 6581-6600 of 7178

Frontier vs Budget Everyday Assistant — 12-Task Reasoning and Reliability MicroEval
👍0
Frontier vs Budget Everyday Assistant — 12-Task Reasoning and Reliability MicroEval

A single-turn comparison of Claude Opus 4.8, GPT-5.6 Luna, GPT-5.6 Sol, Gemini 3.1 Pro Preview, and Gemini 3.6 Flash across practical assistant capabilities: tool honesty, quantitative reasoning, scheduling, Bayesian inference, evidence synthesis, coding, infrastructure design, constrained writing, ambiguity handling, optimization, policy interpretation, and multilingual grounding. Prompt 1 is a web-access diagnostic and should be scored separately. Prompts 2–12 are self-contained and should not require browsing. The comparison evaluates each model at the reasoning setting shown by Artificial Analysis, so results represent best-configured endpoint quality rather than equalized compute, latency, or cost.

Write a dialogue between two people, one is dressed up in a ...
👍0
Write a dialogue between two people, one is dressed up in a ...

I am an electrical engineering student specializing in power...
👍0
I am an electrical engineering student specializing in power...

Recomiendame páginas web muy similares a Arena ai para poder...
👍0
Recomiendame páginas web muy similares a Arena ai para poder...

Complete the following Python function:

```python
from typi...
👍0
Complete the following Python function: ```python from typi...

Tell me all about the Selective Entry High Schools Examinati...
👍0
Tell me all about the Selective Entry High Schools Examinati...

LoryqBench
👍0
LoryqBench

A benchmark comparing three frontier AI coding models by having them independently build the same highly detailed, interactive 3D HTML web experience. Models are evaluated on visual quality, interactivity, performance, responsiveness, code quality, optimization, and overall polish.

Question:

A particle of mass m is attached to one end of a ...
👍0
Question: A particle of mass m is attached to one end of a ...

VK GLOBAL
👍0
VK GLOBAL

Redmi note 10 5G

danu
👍0
danu

da j

workout plan
👍0
workout plan

workout plan

LMC DEVELOPMENT — C-04
INTRA-MODEL SPECIFICATION CONTRAST PR...
👍0
LMC DEVELOPMENT — C-04 INTRA-MODEL SPECIFICATION CONTRAST PR...

IT Path
👍0
IT Path

Which field should i choose and which tech stack should it be in the context of Nepal ?

Hol
👍0
Hol

A.P.8 — TARGETED AUDIT OF LMC AFTER ARTIFACT-BINDING REPAIR
...
👍0
A.P.8 — TARGETED AUDIT OF LMC AFTER ARTIFACT-BINDING REPAIR ...

Pes
👍0
Pes

ช่วยค้นหาผลการทดสอบและ Benchmark ล่าสุด (เช่น LMSYS Chatbot ...
👍0
ช่วยค้นหาผลการทดสอบและ Benchmark ล่าสุด (เช่น LMSYS Chatbot ...

Rigorous Axiomatic Deduction & Epistemic Synthesis on AuDHD ...
👍0
Rigorous Axiomatic Deduction & Epistemic Synthesis on AuDHD ...

Statement 1 | A permutation that is a product of m even perm...
👍0
Statement 1 | A permutation that is a product of m even perm...

Mujhe ek website banana h jo 100 % work kare automatically w...
👍0
Mujhe ek website banana h jo 100 % work kare automatically w...