MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 4381-4400 of 6273

👍0
Hamster Simulation

👍0
Project guidance

👍0
clear

👍0
B.M.9 — FULL METHODOLOGY VALIDATION: v24 ATOMIC DISPATCH-BUN...

👍0
The cyclic subgroup of Z_24 generated by 18 has order
A) 4
...

👍0
Ты — опытный UX/UI-дизайнер и высококлассный conversion-копи...

👍0
qual era il modello migliore con massimo 8B al 13 febbraio?

👍0
Can you code a web service that people will need and like no...

👍0
The Shasu people

👍0
فقد همین تشکر

👍0
المعروف إنه مقاييس الجمال بتتغير مع الوقت
طب كيف تماثيل الإغ...

👍0
Представь себе, магистраль вентиляции самодельная в панельно...

👍0
Estimate the level of recoverable conventional gas reserves ...

👍0
XenonTitan-5526

👍0
Necesito registros bibliograficos (citas bibliográficas) en ...

👍0
Bench on how much model can think for itself.
Almost all the models answer the question with already know facts and details, which being totally wrong, where the correct answer was anything but that generic fact based one.

👍0
Deobfuscation

👍0
arabamı yıkamaya gideceğim ama oto yıkama evime yalnızca 500...

👍0
AWS knowledge test

👍0
Estimate the level and probability of recoverable easily ext...