MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 3241-3260 of 3742

👍0
Which of the following structures travel through the substan...

👍0
Poem

👍0
Actúa como un Ingeniero de Software Senior experto en optimi...

👍0
The cyclic subgroup of Z_24 generated by 18 has order
A) 4
...

👍0
Necip Fazıl Kısakürek'in çile adlı eserinden esinlen erek g...

👍0
Ответ Grok 4.3 (high):
1. Совершенствование системы управлен...

👍0
Just another Local VS Cloud

👍0
Fındıkta bulunan Fito kimyasallar sağlığa faydaları ve mikta...

👍0
Перечисли имена тех писателей, которые писали в стол, не пол...

👍0
DRCC
Email validation

👍0
You are being evaluated on your ability to design and implem...

👍0
==================================================
QISSAH — ...

👍0
Arithmetic
What complexity of arithmetic is it safe to trust LLMs to do without a code interpreter? The purpose of this eval is to understand what level is completely safe, and what level you should instruct and LLM to use a

👍0
Janet’s ducks lay 16 eggs per day. She eats three for breakf...

👍0
sprites

👍0
I want to write an article on dopamine and that how cheap do...

👍0
Research, think, plan, and let me know if it's possible to b...

👍0
Ferry

👍0
Role: You are an expert AI system architect and Python autom...

👍0
zgdocs