MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Multiple-choice questions across 57 academic and professional subjects 

Loading questions…

Public evaluations

Showing 2121-2140 of 3613

BAHAR MAH. 3.BAHÇE ÇIKMA SK. NO: 6 İÇ KApI NO:1 OSMANGAZİ Bu...
👍0
BAHAR MAH. 3.BAHÇE ÇIKMA SK. NO: 6 İÇ KApI NO:1 OSMANGAZİ Bu...

👍0
French correction

Which Gemini Free plan is best at correcting mistakes in French text

prof
👍0
prof

xxx

Eval chat support
👍0
Eval chat support

MoxChatModel

Find the order of the factor group (Z_4 x Z_12)/(<2> x <2>)
...
👍0
Find the order of the factor group (Z_4 x Z_12)/(<2> x <2>) ...

به نظرت دانشگاه درسدن و دانشگاه  اشتوتگارت کدومشون برای علوم...
👍0
به نظرت دانشگاه درسدن و دانشگاه اشتوتگارت کدومشون برای علوم...

You are an administrative operations lead in a government de...
👍0
You are an administrative operations lead in a government de...

generame una pagina en html para una empresa de desarrollo d...
👍0
generame una pagina en html para una empresa de desarrollo d...

sap abap alv report
👍0
sap abap alv report

simple sap abap program with minimal details

Act as a Staff AI Research Engineer and Senior Backend Archi...
👍0
Act as a Staff AI Research Engineer and Senior Backend Archi...

Construct a complete, production-ready software system that ...
👍0
Construct a complete, production-ready software system that ...

test1 of real world prompts for small businesses
👍0
test1 of real world prompts for small businesses

Create a p5.js animation that is a cool interactive thing ba...
👍0
Create a p5.js animation that is a cool interactive thing ba...

hiibb
👍0
hiibb

genara la imagen realista de un perro caminando por la luna
👍0
genara la imagen realista de un perro caminando por la luna

Statement 1 | A permutation that is a product of m even perm...
👍0
Statement 1 | A permutation that is a product of m even perm...

LLM as Designer: Self-Evolving OS Visual Generation
👍0
LLM as Designer: Self-Evolving OS Visual Generation

This benchmark tests an LLM's ability to create a dynamic visual narrative where an AI agent progressively "builds" and "designs" an operating system interface directly in the browser. It combines creative storytelling, dynamic HTML/CSS/JS generation, and animated visualization of a conceptual AI design process. The goal is to make the user feel like they are watching an AI create its own visual environment.

eval
👍0
eval

order
👍0
order

Aungko1990
👍0
Aungko1990

Arfu