MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 8321-8340 of 9810

👍0
guess without context

👍0
Water Simulation: Claude Fable vs Opus

👍0
You are an administrative operations lead in a government de...

👍0
---
15:22
与其说恋爱的隐形作用被可量化的绩点掩盖,不如说社交社会化的重要作用被绩点掩盖了。人是在和社会交互的...
👍0
Mr.
Human
👍0
Black Hole Sim
Black hole simulation single html file

👍0
Pregunta de investigación profunda: ¿Es posible organizar el...
👍0
Estimate the 85th percentile household-networth of the follo...
👍0
Estimate the 85th percentile household-networth of the follo...
👍0
test model jr

👍0
Voxel Kingdom

👍0
You are participating in a staff-level TypeScript technical ...

👍0
test1

👍0
P19.1 — BLINDED EXTERNAL ARCHITECTURE REGRESSION / REPAIR + ...

👍0
jEWß

👍0
Children

👍0
AlphaQuebec-5006

👍0
翻译2023 Acura MDX SH-AWD w/Technology Pkg Sport Utility 4D

👍0
Pregunta de investigación profunda: ¿Es posible organizar el...
👍0
Local LLM Deployment: Intelligence, Hardware, Speed & Cost
Comparison of current open-weight LLMs for fully local/private deployment, evaluating model capability, Q4 memory requirements, inference performance, concurrency, hardware requirements and acquisition cost.