MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 4601-4620 of 7705

👍0
In life, learning what skills or knowledges/subjects is the ...

👍0
tipos de nubes

👍0
SCB SCB-deploy_verify-T1 MUSE

👍0
Guessing output with built-in usage.
This code is trivial to run and see the actual output, but most of the models so far fails spectacularly if they can only guessing and not allows code eval. This test is good to determine if model is actually not bluffing.

👍0
Hazme un generador de tablero de catan

👍0
==================================================
QISSAH — ...

👍0
fastapi test
testing a slm agaist a sota gpt model

👍0
SCB SCB-migration-T1 MK3

👍0
Which LLM modell?

👍0
compara estos dos modelos

👍0
Ok suppose 3.8 gpa at math cs Econ hard courses from ucsd so...

👍0
FIrstEval

👍0
necesito extraer los assets y codigos de un apk, analiza com...

👍0
hffgh

👍0
Obscure Works
Knowledge Depth

👍0
Ideal Agentic Platform
Research to identify a robust solution for accomplishing agentic work in a single interface using capacity from many providers.

👍0
You are a career strategist. This is a single-subject analys...

👍0
Make a 2d mma pvp game that two people can connect online on...
👍0
Напиши 10 очень смешных и крутых шуток, пять с черным юмором...

👍0
Recommend Most AI Agent