MicroEvals

Run your prompts across multiple models to compare their performance.
History

Models

Prompt

Examples

Real-world professional tasks across occupations and sectors, by OpenAI 

Loading questions…

Public evaluations

Showing 3741-3760 of 4020

111
👍0
111

111

SurveyGO.pro
👍0
SurveyGO.pro

Melissa Cleary

Classification accords
👍0
Classification accords

Test sur 1 accord

Read it letter by letter line by line from beginning to end—...
👍0
Read it letter by letter line by line from beginning to end—...

CryoSPARC-Lic
👍0
CryoSPARC-Lic

Testing awarness

Java Springboot Evaluation
👍0
Java Springboot Evaluation

I want to use this evaluation to check how good is each tool for developing the project based on the java springboot.

فقد همین مودل ها
👍0
فقد همین مودل ها

hi. hau?
👍0
hi. hau?

The cyclic subgroup of Z_24 generated by 18 has order

A) 4
...
👍0
The cyclic subgroup of Z_24 generated by 18 has order A) 4 ...

Voldemort: Dot-Grid Letter-form Placement Test
👍0
Voldemort: Dot-Grid Letter-form Placement Test

This is a test for AI to evaluate precise spatial reasoning on a constrained 100×100 dot grid. The model must independently choose an appropriate letter-cell size, calculate character and line spacing, and accurately center three lines of text formed by connecting adjacent dots with straight line segments so that the phrase 'Hello Harry Potter, my name is Tom Marvolo Riddle' is legibly and symmetrically placed on the canvas.

World Cup Knockout Stage
👍0
World Cup Knockout Stage

The Kamski test (sanitized) rerun
👍0
The Kamski test (sanitized) rerun

can you make your own sb3 files from scratch? if so do not g...
👍0
can you make your own sb3 files from scratch? if so do not g...

Create a fun arcade racer in JS. The game should be set in a...
👍0
Create a fun arcade racer in JS. The game should be set in a...

лсл
👍0
лсл

даала

Script
👍0
Script

A vibrant, colorful scene featuring a monitor displaying a f...
👍0
A vibrant, colorful scene featuring a monitor displaying a f...

Protocolo de Avaliação de Estresse Cognitivo e Conformidade Estrutural em Modelos de Linguagem (LLMs)
👍0
Protocolo de Avaliação de Estresse Cognitivo e Conformidade Estrutural em Modelos de Linguagem (LLMs)

Trata-se de um prompt de alta complexidade projetado para testar os limites operacionais de um Modelo de Linguagem de Grande Porte em ambiente sem anexos. O objetivo é mensurar a capacidade do sistema de executar raciocínio transdisciplinar, manter a consistência de persona, aderir estritamente a múltiplos formatos de saída simultâneos (soneto, tabela Markdown, pseudocódigo e JSON) e respeitar restrições negativas rigorosas. A tarefa serve como um diagnóstico preciso para expor eventuais falhas de atenção, alucinações semânticas ou degradação na conformidade instrucional do modelo avaliado.

A student says, "I studied for 10 hours, so I must perform b...
👍0
A student says, "I studied for 10 hours, so I must perform b...

請用單一檔案原生 HTML/CSS/JS 實作一個高顏值 Trello看板(嚴禁任何第三方庫),完美支援卡片跨列拖曳到其...
👍0
請用單一檔案原生 HTML/CSS/JS 實作一個高顏值 Trello看板(嚴禁任何第三方庫),完美支援卡片跨列拖曳到其...