MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 261-280 of 8020

👍1
You are debugging a Linux/Wayland application written in Rus...

👍1
Overall best countries in Africa ranked currently

👍1
Design
Mechnaical pencil

👍1
Easy Problems That LLMs Get Wrong
This MicroEval evaluates LLM responses to simple logic-based questions that LLMs commonly get wrong. The problems used in this MicroEval are from the ArXiv paper of the same name: https://arxiv.org/abs/2405.19616. It is by Sean Williams and James Huckle, so props to them for developing this experiment all the way back in 2024.

👍1
Website Maker

👍1
7x - 9y = 39
35x - 45y = 195
For any real number r, which ...

👍1
curl https://api.avian.io/v1/chat/completions \
-H "Author...

👍1
Hehe
Build it

👍1
yoo

👍1
Confirmación de Directiva ..stbn.
El sistema asimila que el...

👍1
Chess Move Genearation
Test models' ability to count the number of moves in a given chess position

👍1
The Zero-Knowledge Challenge: A Visual Primer (Game)

👍1
Scart light
If you need I will one you Feed

👍1
Bahasa Indonesia
Hanya Keaslian Yang Mampu Mengalahkan Kesem...
THE LAKEFRONT ESTATE MAROS
👍1
keresd meg az összes hibát!
proc parseLegacyNumber(part: str...

👍1
You are an administrative operations lead in a government de...

👍1
Math-Test

👍1
Aoi-chan... ¿has visto a Nene? No la encuentro hace rato. La...

👍1
偶像鸡/idol chicken

👍1
Code me a recreation of the game minecraft for mobile users