Build your own custom benchmark
Standardized benchmarks measure general model capability, but they can't tell you which model is right for your specific use case. Optima lets you build custom benchmarks around your own tasks, so you can compare models on performance, cost, and time efficiency
Benchmark how well models review our supplier contracts — flag risks, extract terms.
Contract Review Benchmark
24 tasks · rubrics attached · 4 categories
Why build your benchmark with Optima
Results specific to your use case
Build a benchmark from your own data and tasks, and get results on the latest models as soon as they're released.
Cut cost and time by over 10x
Compare on more than performance, with every result showing cost and time per task efficiency comparisons across models
Built on Artificial Analysis grading expertise
Grade against custom rubric criteria, or use our panel of judges from major evaluations to rank models head-to-head
How it works
Give Optima your context
Describe the work, attach a few examples, and the build agent drafts the tasks and rubrics with you.
Choose your evaluation type
Example task prompt
Which HS tariff code applies to lithium-ion e-bike batteries?
How it's graded
The rubric carries the expected answer. A judge model checks each response against it, so the score is deterministic — right or wrong, criterion by criterion.
Run across models
Or bring your own agent
Your own agent can compete in the same run, over HTTP.
Grade and decide
| Model | Score | Cost / Task | Time / Task |
|---|---|---|---|
| 60 | $0.18 | 74s | |
| 59 | $0.11 | 52s | |
| 57 | $0.04 | 61s | |
| 50 | $0.02 | 29s |
Pricing
Optima pricing is based on token usage.
- Creating and running a benchmark is charged at the raw token cost of the models used, with nothing added on top.
- Rubric grading is $0.25 per criterion, per model.
- Pairwise grading is $0.75 per match.
At the start of benchmark creation, each benchmark run and each grading pass, we hold an amount of your credit based on our cost estimate for that stage. At the end of the stage you are only ever charged for what was actually used, regardless of the estimate.
Find the best model for your work
Bring your own tasks or describe your use case, and get graded results with the costs attached.
Try Optima