MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 1881-1900 of 3117

👍0
order

👍0
Aungko1990
Arfu

👍0
תיצור לי את כל הקוד לאפליקציה מוכנה שמה שהיא תעשה זה שליטה מ...

👍0
Jarvis UI website test
How does Hallucination affect website development?

👍0
help me learn how EU works from JTI corporate affairs and communication manager perspective, working in italy

👍0
teting
lets see what happen

👍0
hitman codename 47 oyunun karakter modelleri tasarlanırken, ...

👍0
what is span margin

👍0
veste couleur
![[SYSTEM] Convert a natural-language user idea + target aspec...](/_next/image?url=https%3A%2F%2Fartificialanalysiscdn.com%2Fmicro-evals%2F81e2375a0ebb4729b2d1b27869f1bb82.jpg&w=3840&q=75)
👍0
[SYSTEM] Convert a natural-language user idea + target aspec...

👍0
Create a fun arcade racer in JS. The game should be set in a...

👍0
roma

👍0
Book Segmentation
Reasoning ability to see if the model can reason which sentence/text chunk is spoken by which character

👍0
Vì sao nhiều người bảo thủ tính bừa bãi của mình?

👍0
tipos de nubes

👍0
Guessing output with built-in usage.
This code is trivial to run and see the actual output, but most of the models so far fails spectacularly if they can only guessing and not allows code eval. This test is good to determine if model is actually not bluffing.

👍0
==================================================
QISSAH — ...

👍0
Ok suppose 3.8 gpa at math cs Econ hard courses from ucsd so...

👍0
necesito extraer los assets y codigos de un apk, analiza com...

👍0
hffgh