Announcing Optima: create a custom benchmark for your use case
We are launching Optima. Anyone can now create a custom benchmark for their use case, leveraging Artificial Analysis' research and platform: build a benchmark from your own files, agent traces or coding environment, run it across leading models in a single click, and compare quality alongside cost per task and time per task.Build your own benchmark
Introducing Optima
Building and running benchmarks is difficult. We have distilled Artificial Analysis' research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency.
Optima lets you find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task.
How Optima works
We have applied the benchmarking approaches and infrastructure we use at Artificial Analysis across each part of Optima.
- Build benchmarks from your own data and workflows. There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize, Braintrust and Langfuse. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you.
- Run across the latest models. Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released.
- Bring Artificial Analysis grading to your own benchmark. Evaluate responses against objective rubric criteria, or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set.
- Compare performance, cost and time efficiency. Optima benchmarks more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, so you can compare the tradeoffs between models for your specific use case.
What pre-release testers built
Ahead of launch, we gave a group of pre-release testers access to Optima. Examples of benchmarks they created include:
- Which model can save me 10x the cost without a meaningful decrease in quality for my finance and accounting agent?
- Which model best matches the writing style of lawyers for my legal agent?
- Which model can best identify different elements in my custom image dataset?
Get started
Optima is available today. Build your own benchmark, and tag @ArtificialAnlys with what you create.
Read the latest

GPT-6 Sol and Luna push the cost efficiency frontier
GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others
September 22, 2026

Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index
Anthropic's new Opus scores 58 and arrives with a 20% price cut and a larger cache hit discount
September 22, 2026

Benchmarking Grok 4.7
Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs. Coding Agent Index performance has also improved, overtaking GPT-5.6 Sol
September 21, 2026