MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 2861-2880 of 9999

👍0
wres
👍0
Execute tasks as a pure computational synthesis and mathemat...

👍0
HTML WebDev Challenge
A simple web development challenge that tests an LLM's ability to create HTML web pages given a prompt. The eval is 8 prompts. The LLM is ONLY allowed to build with HTML, Javascript, and TailwindCSS. It may pull Javascript libraries from a CDN (like Jsdelivr or Cloudflare), but only if it is SURE they exist and are needed for the build. The first four prompts are very specific, and then the rest give the model more freedom.

👍0
idk
👍0
Hi!
👍0
You are an administrative operations lead in a government de...
👍0
Hola

👍0
What is the best build in mass builder?

👍0
Prove this is true. Do not stop until tou have proved it is ...
👍0
fhn
👍0
calidad de realismo

👍0
You are an administrative operations lead in a government de...

👍0
Create a web-based mobile friendly 3D simulator that allows ...
👍0
ai tools

👍0
Complete the following Python function:
```python
def trunc...
👍0
详细分析 AI Coding 的最佳实践是什么,
目前几个核心文件
1. AGENTS.md
2. PRODUC...

👍0
You are an administrative operations lead in a government de...

👍0
Üst üste dizilmiş bir desimetrekare alanı olan 10 adet alümi...
👍0
qwen 3.8, sonnet 5,
👍0
You are an administrative operations lead in a government de...