MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 5281-5300 of 6045

👍0
«به عنوان تحلیلگر ارشد بازار طلا و ارز ایران عمل کن. با بررس...

👍0
DRCC
Email validation

👍0
You are being evaluated on your ability to design and implem...

👍0
LMC DEVELOPMENT — C-04 THIRD-FAILURE PACKAGING ARCHITECTURE ...

👍0
Hello
👍0
Persona following Spanish

👍0
==================================================
QISSAH — ...

👍0
Your task is to take the provided technical description and ...

👍0
Arithmetic
What complexity of arithmetic is it safe to trust LLMs to do without a code interpreter? The purpose of this eval is to understand what level is completely safe, and what level you should instruct and LLM to use a

👍0
Janet’s ducks lay 16 eggs per day. She eats three for breakf...

👍0
Complete the following Python function:
```python
def how_m...

👍0
You are an AI evaluation engineer specializing in perceptual...

👍0
MASTER BUILD PROMPT — "SELENE PROTOCOL" (working title)
A lu...

👍0
sprites

👍0
Estimate the level of recoverable easily extractable commerc...

👍0
I want to write an article on dopamine and that how cheap do...

👍0
Research, think, plan, and let me know if it's possible to b...

👍0
test1

👍0
Ferry

👍0
instruction following