MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 4801-4820 of 5491

👍0
Arithmetic
What complexity of arithmetic is it safe to trust LLMs to do without a code interpreter? The purpose of this eval is to understand what level is completely safe, and what level you should instruct and LLM to use a

👍0
Janet’s ducks lay 16 eggs per day. She eats three for breakf...

👍0
Complete the following Python function:
```python
def how_m...

👍0
You are an AI evaluation engineer specializing in perceptual...

👍0
MASTER BUILD PROMPT — "SELENE PROTOCOL" (working title)
A lu...

👍0
sprites

👍0
Estimate the level of recoverable easily extractable commerc...

👍0
I want to write an article on dopamine and that how cheap do...

👍0
Research, think, plan, and let me know if it's possible to b...

👍0
test1

👍0
Ferry

👍0
instruction following

👍0
Role: You are an expert AI system architect and Python autom...

👍0
早上起来,先刷牙还是先吃早餐?

👍0
zgdocs

👍0
test add oil

👍0
Janet’s ducks lay 16 eggs per day. She eats three for breakf...

👍0
Best coder

👍0
You are a professional experienced genius intelligent creati...

👍0
What's the most liked video featuring or starring a person w...