MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 2321-2340 of 9102
👍0
Local AI Professional Workload Benchmark
Benchmark designed to compare LLMs for private local deployment serving 1–10 users. It evaluates reasoning, information extraction, instruction following, numerical analysis, prioritization, professional writing, and resistance to unsupported assumptions.
👍0
Testing
Just wanna test AI model
👍0
6Luna - 4.1Flash

👍0
Altın bir ışıkla yıkanır bağlar,
Tahta köprü beni dünüme b...

👍0
szerinted sasha banks vagy bayley ert volna el kevesebbet a ...

👍0
CREA UN JUEGOSIMPLE

👍0
# OUTCOME
Build one complete, production-quality, visually e...
👍0
Real world physics

👍0
fable vs sol
which is best

👍0
Complete the following Python function:
```python
from typi...

👍0
写一段日系青春伤痛文学

👍0
You are a Quantitative Researcher at a proprietary trading f...
👍0
Scene 1 – Village Morning: A beautiful Indian village at sun...

👍0
PG Wodehouse Variations
Shows how well the best models can write.

👍0
Picture
👍0
как различается Breaking Bad и Better Call Saul? в чем главн...

👍0
دو فایل دارم آنها را اجرا کن

👍0
I am reverse-engineering a synthetic OTC binary options pric...
👍0
Test
👍0
Una mujer en bikini roja muy diminuta y con pechos muy grand...