MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 3141-3160 of 6558

👍0
5.5pro

👍0
Where does your favourite LLM rank currently?

👍0
Дарова

👍0
You are an administrative operations lead in a government de...

👍0
Compare all reasoning effort levels in GPT-5.6 Luna.

👍0
Viral LLM Coding Challenges

👍0
Основываясь на предоставленных данных из файлов, я провел де...

👍0
Aldrich Ocampo Nuevo aldrichnuevo369@gmail.com

👍0
Estimate the level of recoverable easily extractable commerc...

👍0
ki trading
strict prompt following

👍0
A robe takes 2 bolts of blue fiber and half that much white ...

👍0
P22.5 — INTERNAL SHADOW / WEB RADAR / INTER-RESULT WORK / AG...

👍0
egy nema refluxosnak lehet halalos vagy extra veszelyes egy ...

👍0
web search ability
web search ability

👍0
i wanna say to my Moroccan friend in whatssap to come visit ...

👍0
Jsi nezávislý red-team evaluátor řídicího promptu LLM. Odpov...

👍0
strategy comparison
evaluate how good is my commercial approach

👍0
Complete the following Python function:
```python
def car_r...

👍0
Создай премиальный, современный, высококонверсионный одностр...

👍0
galaxy