MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 241-260 of 7822

👍1
Perplexity AI:
Perplexity AI:

👍1
GAS Örnek İhtiyacı
👍1
Herra Agentti
Asshole

👍1
quiero que me transfieras a código termux infinito con esta ...

👍1
"This is a small puppy with light yellow fur, a cute face, a...

👍1
SDF Creation

👍1
Pudding
physics test

👍1
Test for reasoning

👍1
IQ test Generator
An visual IQ test generator

👍1
devcontainer

👍1
vbyes
vb
👍1
Life?

👍1
3d state of the art raisoning

👍1
Startup Idea Brainstorming

👍1
体育赛事网站
需要AI提供一个较为完整的体育赛事管理网站,包括前后端,涵盖小组赛和淘汰赛的一系列功能。

👍1
You are debugging a Linux/Wayland application written in Rus...

👍1
Overall best countries in Africa ranked currently

👍1
Design
Mechnaical pencil

👍1
Easy Problems That LLMs Get Wrong
This MicroEval evaluates LLM responses to simple logic-based questions that LLMs commonly get wrong. The problems used in this MicroEval are from the ArXiv paper of the same name: https://arxiv.org/abs/2405.19616. It is by Sean Williams and James Huckle, so props to them for developing this experiment all the way back in 2024.

👍1
Website Maker