MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 1161-1180 of 6186

👍0
Dame los planos para crear una maquina vibratoria de mano ej...

👍0
Actúa como un asistente organizador y diseñador de automatiz...

👍0
MMMU、GAIA、HLE

👍0
Как писали и что писали святые, богословы и учителя церкви п...

👍0
English web
English teaching app

👍0
ငါ့ကို cinematic video အမျိုးစား ဘောလိ၀ုဒ် အက်ရှင်ဇာတ်ကြမ်းက...

👍0
=== HARDWARE GROUND TRUTH: 2x NVIDIA Tesla T4 (Kaggle) ===
A...

👍0
LMC DEVELOPMENT — INDEPENDENT METHODOLOGY AUDIT
ÚKOL
Prove...

👍0
Quiero que uses tu esfuerzo al máximo para que me enseñes mu...

👍0
Write a function to sort a given matrix in ascending order a...

👍0
Benchmark mature, production-ready AI models for IntelFactor...

👍0
Lo que mide exactamente esta prueba al enviarla a una IA:
Pr...

👍0
All Purpose
All purpose tests

👍0
Genaro villegas check
👍0
eval-1788238658571

👍0
prepare a .conf to .lua migration plan. help me master lua. ...

👍0
The Surrealist Dream Logic Room Challenge
This benchmark tests an LLM's ultimate zero-shot creativity and advanced coding skills. The goal is to evaluate its ability to interpret abstract, surreal, and metaphorical concepts and translate them into a coherent, interactive 3D experience using Three.js. This goes beyond standard code generation to test true conceptual modeling.

👍0
If my future wife has the same first name as the 15th first ...

👍0
You are the read-only AI supervisor of a heterogeneous enter...

👍0
Graphing studio
Math graphing studio