MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 6201-6220 of 8798

👍0
以下是融合优化后的提示词: 请依据《红楼梦》原著对大观园的描写,以微体素(细小颗粒体素)风格搭建一个可交互的三维大观园场...

👍0
Hamster Simulation

👍0
Project guidance

👍0
clear

👍0
B.M.9 — FULL METHODOLOGY VALIDATION: v24 ATOMIC DISPATCH-BUN...

👍0
The cyclic subgroup of Z_24 generated by 18 has order
A) 4
...
👍0
A type of small mammal from the mountain regions of the west...

👍0
Ты — опытный UX/UI-дизайнер и высококлассный conversion-копи...

👍0
qual era il modello migliore con massimo 8B al 13 febbraio?

👍0
You are "CryptoGuard," an autonomous, long-running cryptocur...

👍0
Can you code a web service that people will need and like no...

👍0
The Shasu people

👍0
فقد همین تشکر

👍0
المعروف إنه مقاييس الجمال بتتغير مع الوقت
طب كيف تماثيل الإغ...
👍0
P46.1 — HUMAN-USE NONQUALIFIED-STATE / ACTIONABILITY / TRANS...

👍0
Представь себе, магистраль вентиляции самодельная в панельно...

👍0
Estimate the level of recoverable conventional gas reserves ...

👍0
XenonTitan-5526

👍0
Necesito registros bibliograficos (citas bibliográficas) en ...

👍0
Bench on how much model can think for itself.
Almost all the models answer the question with already know facts and details, which being totally wrong, where the correct answer was anything but that generic fact based one.