MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 4141-4160 of 5201

👍0
Top-down luxury cosmetic photography featuring Fino Premium ...

👍0
The Impossible Object Sculpture Park
This benchmark tests an LLM's ability to translate paradoxical, metaphorical, and logically inconsistent concepts into a functional and interactive 3D scene using Three.js. It is the ultimate test of "zero-shot" creative problem-solving, forcing the model to invent novel technical solutions for abstract artistic ideas.

👍0
合成生物學

👍0
Define
\[p = \sum_{k = 1}^\infty \frac{1}{k^2} \quad \text{a...

👍0
Game
Test

👍0
Prosper

👍0
Tifa, yuna from ffx and aerith all as mizutsune. They still ...

👍0
Spx
Spx changing

👍0
REAL USAGE
Only real usage

👍0
math reasoning

👍0
test

👍0
Test1

👍0
You are an administrative operations lead in a government de...

👍0
# 全面实施提示:50亿参数混合稀疏MoE语言模型训练系统
构建一个完整的生产级训练系统,用于训练一个50亿参数混合稀...

👍0
Hi你好

👍0
Complete the following Python function:
```python
from typi...

👍0
kişniş tohumunu n bağırsak sağlığı üzerinde ki etkileri

👍0
test

👍0
Jsi expertní vědecký, medicínský a Evidence-Based Medicine a...

👍0
Testing