MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 401-420 of 3472

👍0
aa_tTdpgMZIsKIGsjIuECgxitkMcjbVBRLZ

👍0
search

👍0
create a video where cosmic spider man is fighting with goku...

👍0
chatgpt vs cloude

👍0
Bbb
Bbnnn

👍0
Statement 1 | If f is a homomorphism from G to K and H is no...

👍0
大型语言模型(LLMs)已日益被部署为能够在广泛任务中执行规划、工具使用和多步推理的自主代理。近期的基于代理的系统在科学...

👍0
i want h2 mach iv 750, find out the best h2 mach iv from the...

👍0
FalconPapa-8889

👍0
translate the text to full english. do not change, edit, sho...

👍0
Unreal Engine 4 Wheel Motion Blur

👍0
5+5

👍0
Test

👍0
Какой был издан на русском языке самый огромный по страницам...

👍0
Regarding your question, Bain and McKinsey do not really do ...

👍0
is it joever for swe
r we cooked

👍0
एक छोटे से गाँव में रामू नाम का गरीब किसान रहता था। वह सुबह ...

👍0
I have read Enoch, BOOK OF JUBILEES, Gospel of Thomas, Lette...

👍0
Practical Coding Reasoning
A focused set of three coding tasks designed to test practical reasoning and implementation skills in real-world scenarios: robust string parsing with edge cases, async concurrency control with ordered results, and immutable nested state updates. Each prompt omits the explicit answer, making it suitable for evaluating both problem understanding and solution correctness without leakage of target outputs.

👍0
MuseSpark1.2
t2