MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 621-640 of 3052

👍0
MMMU、GAIA、HLE

👍0
Как писали и что писали святые, богословы и учителя церкви п...

👍0
Write a function to sort a given matrix in ascending order a...

👍0
All Purpose
All purpose tests

👍0
Genaro villegas check

👍0
prepare a .conf to .lua migration plan. help me master lua. ...

👍0
The Surrealist Dream Logic Room Challenge
This benchmark tests an LLM's ultimate zero-shot creativity and advanced coding skills. The goal is to evaluate its ability to interpret abstract, surreal, and metaphorical concepts and translate them into a coherent, interactive 3D experience using Three.js. This goes beyond standard code generation to test true conceptual modeling.

👍0
If my future wife has the same first name as the 15th first ...

👍0
You are the read-only AI supervisor of a heterogeneous enter...

👍0
Continuous Self Update and Self Improvement

👍0
Video
This Is NOT Normal in America

👍0
Analyst Benchmark

👍0
Tell me about the EMM Magic Wand

👍0
zigl

👍0
je veux obtenir un carnet en .pdf qui me permet de connaître...

👍0
Use the provided master reference sheet as the only visual r...

👍0
HC-BND

👍0
Qual seu modelo?

👍0
SNAP_VGROUP = "Snap3"
DIST_VGROUP = "DistZone3"
def sele...

👍0
3D Neon Cyberpunk HTML