Optima

Build your own custom benchmark

Standardized benchmarks measure general model capability, but they can't tell you which model is right for your specific use case. Optima lets you build custom benchmarks around your own tasks, so you can compare models on performance, cost, and time efficiency

Contract Review Benchmark
Build agent drafting tasks

Benchmark how well models review our supplier contracts — flag risks, extract terms.

Working
Reading your example contracts
Drafting grading rubrics
Writing task 18 of 24

Contract Review Benchmark

24 tasks · rubrics attached · 4 categories

Why build your benchmark with Optima

Results specific to your use case

Build a benchmark from your own data and tasks, and get results on the latest models as soon as they're released.

Cut cost and time by over 10x

Compare on more than performance, with every result showing cost and time per task efficiency comparisons across models

Built on Artificial Analysis grading expertise

Grade against custom rubric criteria, or use our panel of judges from major evaluations to rank models head-to-head

How it works

1

Give Optima your context

Describe the work, attach a few examples, and the build agent drafts the tasks and rubrics with you.

Describe your use case|Attach examplespptxAA-Merch_Sales_Pitch_DeckpptxAA-Merch_Startup_Swag_Pitchtwo decks we were happy with — this is what good looks likeBuild agentdrafts tasks + rubricsYour drafted benchmarkSales Deck Creation5 tasks · 10 criteria · head-to-headNorthstar Hotels welcome-kit pilotCrescent Arts Museum collection pitchApex Trails 12-month merch programmeTidepool summer pre-season launchCity Sound Festival activation
2

Choose your evaluation type

Task style
Objective
Rubric judge
Subjective
Optima
Q&A
The model answers questions that have known correct answers
Document input
The model answers questions about your uploaded files
Agentic
The model completes tasks and produces deliverables, using tools along the way
Interaction
Simulate a real conversation with different types of users
Coming soon
Q&AObjective grading

Example task prompt

Which HS tariff code applies to lithium-ion e-bike batteries?

How it's graded

The rubric carries the expected answer. A judge model checks each response against it, so the score is deterministic — right or wrong, criterion by criterion.

3

Run across models

Your benchmark24 tasksFlag the risky clausesExtract the payment termsDraft the counterparty summarysame tasks, same conditions, every modelRunsandbox + toolsModels you pickedClaude Fable 5GPT-5.6 SolKimi K3Gemini 3.6 Flashtokens, cost and time recorded per task24/24

Or bring your own agent

Your own agent can compete in the same run, over HTTP.

Claude Opus 4.6GPT-5.2Gemini 3 Proyour-agentPOST /runs/9f2c/artifactsBenchmarkrun24 tasks · same judgesLeaderboard1Claude Opus 4.63GPT-5.24Gemini 3 Pro2your-agent
4

Grade and decide

Contract Review Benchmark24 tasks · 4 models
ModelScoreCost / TaskTime / Task
Claude Fable 560$0.1874s
GPT-5.6 Sol59$0.1152s
Kimi K357$0.0461s
Gemini 3.6 Flash50$0.0229s
strong and cheapScoreCost per task$0.00$0.05$0.10$0.15$0.20Claude Fable 560 · 74s per taskGPT-5.6 Sol59 · 52s per taskKimi K357 · 61s per taskGemini 3.6 Flash50 · 29s per task

Pricing

Optima pricing is based on token usage.

  • Creating and running a benchmark is charged at the raw token cost of the models used, with nothing added on top.
  • Rubric grading is $0.25 per criterion, per model.
  • Pairwise grading is $0.75 per match.

At the start of benchmark creation, each benchmark run and each grading pass, we hold an amount of your credit based on our cost estimate for that stage. At the end of the stage you are only ever charged for what was actually used, regardless of the estimate.

Find the best model for your work

Bring your own tasks or describe your use case, and get graded results with the costs attached.

Try Optima