MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Multiple-choice questions across 57 academic and professional subjects 

Loading questions…

Public evaluations

Showing 2721-2740 of 2992

Hola
👍0
Hola

یک برنامه بساز که کار مفید انجام بده  و کارهای روزمره رو راح...
👍0
یک برنامه بساز که کار مفید انجام بده و کارهای روزمره رو راح...

Okjia
👍0
Okjia

Ok

joever for soydev or not
👍0
joever for soydev or not

hey chat r we cooked

Find the 2024 state of texas ISD robinhood recapture amounts...
👍0
Find the 2024 state of texas ISD robinhood recapture amounts...

MicroEval 1
👍0
MicroEval 1

Render a rotating kebab in javascript.
👍0
Render a rotating kebab in javascript.

ازت می‌خوام به عنوان یک کتابخوان حرفه ای بهترین ترجمه کتاب 1...
👍0
ازت می‌خوام به عنوان یک کتابخوان حرفه ای بهترین ترجمه کتاب 1...

fait un jeu: models non simplifiés, le jeu à deux IA une typ...
👍0
fait un jeu: models non simplifiés, le jeu à deux IA une typ...

Ein niedlicher grüner Frosch mit blauem Hoodie hält einen kl...
👍0
Ein niedlicher grüner Frosch mit blauem Hoodie hält einen kl...

Ankara yaris programini analiz yapip sansli atlari bul
👍0
Ankara yaris programini analiz yapip sansli atlari bul

Das Ost ein Tets
👍0
Das Ost ein Tets

hy how are you
👍0
hy how are you

Photo empathique et authentique : jeune maman algérienne ass...
👍0
Photo empathique et authentique : jeune maman algérienne ass...

Minecraft 001
👍0
Minecraft 001

001

General Purpose Prompts
👍0
General Purpose Prompts

Make a cat logo
👍0
Make a cat logo

Frontier vs Budget Everyday Assistant — 12-Task Reasoning and Reliability MicroEval
👍0
Frontier vs Budget Everyday Assistant — 12-Task Reasoning and Reliability MicroEval

A single-turn comparison of Claude Opus 4.8, GPT-5.6 Luna, GPT-5.6 Sol, Gemini 3.1 Pro Preview, and Gemini 3.6 Flash across practical assistant capabilities: tool honesty, quantitative reasoning, scheduling, Bayesian inference, evidence synthesis, coding, infrastructure design, constrained writing, ambiguity handling, optimization, policy interpretation, and multilingual grounding. Prompt 1 is a web-access diagnostic and should be scored separately. Prompts 2–12 are self-contained and should not require browsing. The comparison evaluates each model at the reasoning setting shown by Artificial Analysis, so results represent best-configured endpoint quality rather than equalized compute, latency, or cost.

Write a dialogue between two people, one is dressed up in a ...
👍0
Write a dialogue between two people, one is dressed up in a ...

I am an electrical engineering student specializing in power...
👍0
I am an electrical engineering student specializing in power...