MicroEvals
Run your prompts across multiple models to compare their performance.
Public evaluations
Showing 1081-1100 of 3596

👍0
"""
AgentOS – TaskPlan ORM model.
Stores the selected execu...

👍0
Complex Scenario Solution Coder

👍0
The Most AI Agent

👍0
SVG art test
LLM SVG interpretability

👍0
Act as a Principal UI/UX Designer and Senior Frontend Engine...

👍0
Workflow
engineering

👍0
creator power

👍0
Cartoon video
Cartoon video

👍0
Riddle

👍0
test

👍0
The mass of Saturn's rings is 2x1019 kg. What is the ratio o...

👍0
Testing comisioning tecnical
Server

👍0
bvgcg
nbhvvg

👍0
1+8888888等于几

👍0
An enormous alpine lake at midnight, perfectly still water r...

👍0
hãy tìm cachs để dùng https://artificialanalysis.ai/ api để ...

👍0
Make a roadmap with the best resources to learn CS + Mathema...

👍0
price and cost

👍0
DBD-Benchmark
A benchmark made specifically to test LLMs with common sense questions. Based on my alredy existing eval on huggingface: Martico2432/DBD-Benchmark

👍0
移动同步
移动对鼠标移动轨迹进行算法修正,提高追踪的稳定性
直线修正
在鼠标直线移动时修正轨迹偏差,清除偏移抖动
波浪修...