MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 3821-3840 of 6424

👍0
Book Segmentation
Reasoning ability to see if the model can reason which sentence/text chunk is spoken by which character

👍0
HumanEval Question 32
HumanEval Question 32

👍0
设计飞机空调
空调设计计划

👍0
Vì sao nhiều người bảo thủ tính bừa bãi của mình?

👍0
In life, learning what skills or knowledges/subjects is the ...

👍0
tipos de nubes

👍0
Guessing output with built-in usage.
This code is trivial to run and see the actual output, but most of the models so far fails spectacularly if they can only guessing and not allows code eval. This test is good to determine if model is actually not bluffing.

👍0
Hazme un generador de tablero de catan

👍0
==================================================
QISSAH — ...

👍0
In 2015, he said, a study completed in cooperation with the ...

👍0
Which LLM modell?

👍0
compara estos dos modelos

👍0
Ok suppose 3.8 gpa at math cs Econ hard courses from ucsd so...

👍0
FIrstEval

👍0
necesito extraer los assets y codigos de un apk, analiza com...

👍0
hffgh

👍0
Obscure Works
Knowledge Depth

👍0
You are a career strategist. This is a single-subject analys...
👍0
Напиши 10 очень смешных и крутых шуток, пять с черным юмором...

👍0
Recommend Most AI Agent