MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 2361-2380 of 3949

👍0
tipos de nubes

👍0
Guessing output with built-in usage.
This code is trivial to run and see the actual output, but most of the models so far fails spectacularly if they can only guessing and not allows code eval. This test is good to determine if model is actually not bluffing.

👍0
==================================================
QISSAH — ...

👍0
Ok suppose 3.8 gpa at math cs Econ hard courses from ucsd so...

👍0
necesito extraer los assets y codigos de un apk, analiza com...

👍0
hffgh

👍0
Recommend Most AI Agent

👍0
First person who discovered that Earth is round.

👍0
How is the optimal automatic bot trading setup that have the...

👍0
test

👍0
Toppers Game Developing
Game Developing

👍0
You’ve been hired as an In Ear Monitor (IEM) Tech for a tour...

👍0
hub

👍0
compare top models for coding, but include rating by output ...

👍0
UNIVERZÁLNÍ ONE-SHOT DECISION-THEORY STRESS TEST 4
ÚČEL
Jsi...

👍0
Ok suppose 3.8 gpa at math cs Econ hard courses from ucsd so...

👍0
একটা স্ক্রিপ্ট যেটা সুন্দর করে আমার ডাটাবেসের পুরা স্ট্রাকচা...

👍0
Idk
Idk

👍0
Crie um site para uma agência de design gráfico e desenvolvi...

👍0
IQ Test Generator
Creates dynamic IQ Tests