MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 1601-1620 of 3649

👍0
The cyclic subgroup of Z_24 generated by 18 has order
A) 4
...

👍0
123

👍0
Classic arcade games

👍0
can you write a sentence that uses an em dash

👍0
fait un simulateur de survie en mer sur un bâteau a moteur à...

👍0
Complete the following Python function:
```python
from typi...

👍0
Find the order of the factor group (Z_4 x Z_12)/(<2> x <2>)
...

👍0
You are a highly specialized DepEd MATATAG Key Stage 3 (Grad...

👍0
Silly token in/out benchmark (short)
How fast does each model read a long prompt and print out a long answer. Measures raw time-to-solution for a predictable, fixed output. Short variant.

👍0
analyze my resume - Sahil Bhatt + Toronto, ON # sahil.bhatt@...

👍0
"""Host a downloaded JAIDE checkpoint as a Modal web service...

👍0
Which parameter has the greatest impact on how quickly a tas...

👍0
ptvgff

👍0
Consider a universe consisting of a vast network of nodes (r...

👍0
Best Winning Project developing, writing and producing model

👍0
Output of ('hhh'*0) in python and why is the output is given

👍0
zelda breath of the wild
Three.js game offline capabilities

👍0
Pipe Puzzle
classic game; by mnf

👍0
JVM CORE VHDL

👍0
simple reasoning test
simple reasoning test that trips up older or smaller models set to low reasoning