MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 3881-3900 of 8984

👍0
math
math
👍0
sdasdsada

👍0
Silly token in/out benchmark (short)
How fast does each model read a long prompt and print out a long answer. Measures raw time-to-solution for a predictable, fixed output. Short variant.

👍0
analyze my resume - Sahil Bhatt + Toronto, ON # sahil.bhatt@...

👍0
Complete the following Python function:
```python
from typi...

👍0
chem ds
👍0
Create a single-file index.html using Three.js with OrbitCon...

👍0
"""Host a downloaded JAIDE checkpoint as a Modal web service...
👍0
Create a single-file index.html using Three.js with OrbitCon...

👍0
Make an interactive education app UI like Duolingo, Brillian...

👍0
strarberry有几个r
strarberry有几个r

👍0
Estimate the 85th percentile household net worth of the foll...

👍0
Which parameter has the greatest impact on how quickly a tas...

👍0
Experiments
123

👍0
How to learn mathematics well
👍0
CSV adjuster
👍0
Trabalho de inglês
7 ano 1

👍0
ptvgff

👍0
Consider a universe consisting of a vast network of nodes (r...

👍0
Best Winning Project developing, writing and producing model