Image Generation Benchmarking Methodology

Scope & Background

Artificial Analysis benchmarks image generation models typically delivered via serverless API endpoints. This page describes how we measure their quality, speed, and price.

We cover two image generation modalities:

  • Text to Image: models that generate an image from a text prompt only.
  • Image Editing: models that modify a reference image based on a text prompt.

Endpoint Coverage

Covered workloads: inference benchmarking covers Text to Image generation endpoints.

Public serverless endpoints only: benchmarking covers true serverless endpoints that are, or will be, publicly available. Demo products, hardware showcases, and dedicated or private deployments are out of scope.

Standard & Modified Endpoints

We categorize each endpoint as either Standard (serving the native model at full fidelity, with its standard input parameters exposed) or Modified (altered from the native model in ways that affect output fidelity, typically to increase generation speed or reduce cost).

Modified endpoints are labeled as Modified on Artificial Analysis and may be hidden from the default leaderboard view, remaining visible to users who opt to display them. To preserve the fairness of our default benchmarks, we retain discretion over whether an endpoint is categorized as Modified — indicators include distillation, partial exposure of the native model's input parameters, or materially degraded output quality — and over how many variants of a single model are listed.

Default Generation Settings

We generate images using each model's published defaults. For first-party APIs, this means the API's documented defaults; for open-source models, the defaults from the open-source repository. This covers inference steps, guidance scale, and any default negative prompt behavior. For fair comparison across models, we apply the following normalizations:

  • We generate at the highest resolution each model supports, then downscale to 1024×1024 to serve in our Arenas.
  • We generate 1 image per prompt.
  • We use a 1:1 aspect ratio.
  • We use a seed of 42.

For Image Editing models, each prompt is paired with a reference image from a curated set.

Key Metrics

We use the following metrics to track quality, performance, and price for image generation models.

Quality Elo

Each model's Elo score reflects its relative quality, derived from user votes in the Image Arena. We compute ratings using Bradley-Terry Maximum Likelihood Estimation and rescale them to an Elo-like range for readability.

Each modality is scored independently, as Arena matchups only pair outputs from the same modality.

Price per 1,000 Images

Provider's price (USD) per image generated, multiplied by 1,000.

Pricing is taken directly from each provider's published per-image rate.

Generation Time

Median time the provider takes to generate a single image, calculated over the past 3 days.

Images are generated at a batch size of 1. Generation Time includes downloading the image from the provider where a URL is provided rather than an image response. This reflects end-user latency, since URLs can be issued before generation completes.

Text to Image

Methodology Overview

Text to image models vary in capability across use cases. One model may lead on text rendering, another on complex layouts or photorealistic human characters. The model an agency chooses for campaign creative is not necessarily the one a product manager chooses for UI mockups, or a game studio for concept art.

Our methodology measures this variation directly. We rank models on human preference across a taxonomy of real-world use cases and model capabilities, identifying the best model overall and the best model for specific work. The taxonomy anticipates where the field is going, covering the capabilities leading labs are pursuing and the use cases where image generation is seeing real adoption. We refresh our prompt set monthly so the leaderboard continues to measure the frontier as it moves.

Prompt Curation and Refresh

We maintain a rotation of prompts, refreshed regularly to track the frontier as model capabilities evolve and to keep our leaderboard fair and accurate.

  • Prompt Taxonomy: Every prompt is authored against a taxonomy of two axes, real-world use case and model capability, and tagged for both. Prompts are sampled evenly across the taxonomy in our arena.
  • Real User Data: Our prompts are written to reflect how end users actually prompt, informed by anonymized crowdsourced data and a human-curated, live-updated prompt corpus.
  • Freshness: We refresh our prompt set every month. Prompts are retired based on how well they discriminate between models and how well they reflect the way real users prompt image models today.

Our prompt set carries the signal of real user data while providing unbiased, even coverage across use cases and capabilities.

Prompt Taxonomy

We construct our prompt taxonomy along two axes.

Use Case. These use cases are grounded in our analysis of how consumers and enterprises are adopting text to image generation, and our observation of emerging use cases. Example tasks for each include:

  • Marketing & Advertising: Ad creatives, campaign visuals, posters, brand imagery
  • Retail & E-commerce: Product shots, on-model apparel, packaging mockups
  • Live-Action Film: Cinematic scenes, film stills, storyboards, character sheets
  • Animation & Gaming: Game assets, concept art, character design, comics
  • Architecture & Real Estate: Interior and exterior visualization, floor plans, staging
  • Productivity & Knowledge Work: Diagrams, infographics, charts, slides
  • UI/UX Design: UI mockups across app, web, in-car, and spatial surfaces
  • Consumer: Book covers, stock images, editorial illustration
  • Social Media & Creator Content: Thumbnails, promo cards, channel and profile art
  • Frontier: Capabilities at and beyond the current frontier

Capability. These model capabilities are informed by our ongoing work with leading model labs and the capabilities they are pursuing, and grounded in academic benchmarks. Examples for each include:

  • Reasoning: Entity, mathematical, spatial, logical reasoning, concept mixing, idiom interpretation
  • Knowledge: Real landmarks, species, and domain facts across science and common sense
  • Text Rendering: Long text, small text, symbols, artistic lettering
  • Layout: Flows, arrows, blocks, visual hierarchy, multi-panel compositions
  • Complex Compositions: Precise counting, spatial relationships, subject interactions, attribute binding
  • Lighting: Reflection, refraction, shadows, caustics
  • Material: Surface properties, transparency, subsurface scattering, texture realism
  • Physics: Gravity, support, collision, thermal and state change
  • Human Anatomy: Hands, faces, body proportion, dynamic poses and movement

Prompt Writing

Prompts follow strict authoring standards: plain natural language, a single positive prompt with the negative-prompt field left empty, and architecture-neutral phrasing that does not favor any model family. Prompts are screened for redundancy at authoring time so they cover distinct real-world usage scenarios. Prompts are written:

  • For full coverage of the taxonomy, distributed evenly across every use case and capability combination. Because each prompt carries a use case, capability, and style tag, we can evaluate both overall model quality and performance on the specific cuts that matter to each user. One user needs the best model for UI design with complex layout hierarchy, another might need cinematic shots with realistic human characters.
  • To reflect how end users really prompt, informed by crowdsourced anonymized prompts from real users and a human-curated, live-updated prompt corpus of consumer prompting patterns.

All prompts are:

  • English language
  • Human curated

This produces high-signal prompt sets that evaluate models against the generation tasks most relevant to end users and industry, today and in the near future.

Prompt Retirement

Every month, we select prompts to retire based on two criteria:

  • Discrimination: whether each prompt still produces clear signal about model capability differences. For each vote, we label the higher-rated model (per the global leaderboard) the favorite. Prompts whose favorite win rate is statistically indistinguishable from chance are flagged for retirement.
  • Freshness/Realism: we compare our benchmark prompt set against our live prompt corpus, including our latest crowdsourced prompts and human-curated prompt corpus, to filter out prompts that no longer reflect real-world prompting conventions.

Retired prompts are no longer served in the arena for vote collection. A monthly refresh cadence also protects the integrity of the leaderboard: a prompt set that rotates cannot be overfit.

Sampling Methodology

Each prompt is tagged with 1 use case and 1 primary capability tag. We sample prompts evenly across each use case and capability combination.

Every matchup pairs two outputs generated from the same prompt, within the same modality. Left/right position is randomized and model identities are revealed only after the vote. Newly added models are temporarily upsampled until their ratings converge, and opponents are drawn to maximize the information value of each matchup. Each use case and capability contributes equal weight to the overall leaderboard rating.

Elo Calculation

Vote Collection

Model quality is measured by human preference.

  • Blind pairwise voting: Evaluators see two outputs generated from the same prompt by two different models, without model identities, and select the one they prefer.
  • Judging hints: Each matchup surfaces short hints tied to the prompt's use case, capability, and style tags, directing attention to the specific aspects under test on complex prompts.
  • Engagement gate: Votes can only be cast after a minimum engagement time with each output.
  • Vote quality: All votes pass through bot detection and anomaly filtering before entering the rating calculation.

Vote Filtering

We filter votes according to what our current methodology measures. Votes that predate the current methodology were cast on prompts written to separate an earlier generation of models. Current models saturate many of those prompts, with nearly all producing an acceptable result, so a vote on them records a coin flip rather than a capability difference. We therefore apply a cohort-based filter to historical votes:

  • Current cohort: Models that are publicly accessible, recently released, and most relevant to our audience. Ranked on votes collected under the current methodology only.
  • Legacy cohort: All other models. Ranked on every vote they appear in, preserving their representation on the leaderboard.
  • Rule: A vote is retained if and only if both participating models are eligible for it.

We apply this filter after our standard vote-quality filters, including bot and spam exclusion. Elo is recalculated from scratch over the entire retained vote set, with every model, current cohort and legacy alike, ranked in that single calculation. The rating is anchored on FLUX.1 [schnell] = 1000 for the overall leaderboard, and FLUX.2 [dev] = 1000 for all subcategory leaderboards.

Image Editing

Methodology Overview

At Artificial Analysis, we treat Image Editing and Reference to Image as two independent benchmarks, based on where each capability sits in end-user workflows.

For Image Editing, we test a model's ability to make specific, localized changes to an image while keeping the rest of the image unchanged. We see this benchmark as covering the “post-production” stage of image creation workflows.

For our upcoming Reference to Image benchmark, we test a model's ability to generate entirely novel images that retain characteristics of the input images as a control point. We view this capability as similar to text to image, in the sense that it produces novel outputs rather than modifying existing ones.

For Image Editing, we rank models on human preference across a structured taxonomy of real-world use cases and editing actions, identifying both the best model overall and the best model for specific editing work. Different use cases produce very different editing requests: reliably reshooting scenes for an AI video project is a different problem from altering graphic design for marketing assets, and the best model for enhancement and restoration might not be the best for text and symbol editing.

Prompt Curation and Refresh

Every item in our prompt set is a curated pair: a source image and an edit instruction written against it.

  • Source Images: A curated set of both real and synthetic images.
  • Prompt Taxonomy: Every image editing prompt is authored against two axes, real-world use case and editing action, and tagged for both, following the same methodology as Text to Image. Prompts are sampled evenly across the taxonomy in our arena.
  • Real User Data: Our prompts are written to reflect how end users actually prompt, informed by anonymized crowdsourced data and a human-curated, live-updated prompt corpus.
  • Freshness: We refresh our prompt set every month, following the same methodology as Text to Image.

Prompt Taxonomy

Use Case. The same ten use cases we measure for Text to Image: Marketing & Advertising, Retail & E-commerce, Live-Action Film, Animation & Gaming, Architecture & Real Estate, Productivity & Knowledge Work, UI/UX Design, Consumer, Social Media & Creator Content, and Frontier.

Editing Action. What the editing prompt asks the model to do:

  • Object-Level Edit: Adding, removing, replacing, or altering specific objects in the frame
  • Scene & Style Edit: Relighting, restyling, and changes to the background, setting, or overall treatment
  • Enhancement & Restoration: Denoising, retouching, artifact removal, and repair of damaged or degraded images
  • Composition & Framing: Cropping, extending, reframing, and changes to camera view or perspective
  • Identity-Preserving Edit: Changes to people, faces, and characters that must keep identity intact
  • Text or Symbol Edits: Editing text, typography, logos, and symbols within the image
  • Reasoning-Based Edit: Edits requiring inference about the scene rather than a literal instruction

Elo Calculation

Elo calculation follows the same methodology as Text to Image, including vote filtering: models in the current Image Editing cohort are ranked only on votes collected under the current prompt set, while legacy models keep their full vote history, and a vote is retained if and only if both participating models are eligible for it. The rating is anchored on Gemini 2.0 Flash Preview = 1000 for the overall leaderboard, and FLUX.2 [dev] = 1000 for all subcategory leaderboards.

Generation Time Testing Methodology

Key technical details:

  • Cadence: benchmarks run every 2 hours at a random offset within each window.
  • Sample size: the headline Generation Time is the median of successful measurements from the trailing 3 days, typically around 36 samples per host model per window.
  • What is included: each measurement times the full request lifecycle for one image, covering the API request, provider inference, and either the inline response body or the URL download to disk. Where providers return a URL, the URL is often emitted before the image is fully written, so downloading the bytes is counted to reflect real end-user latency.
  • Unique prompts: a new prompt is generated for each call from a curated pool, so providers cannot return cached responses across runs.
  • Resolution: we generate at 1024×1024. If a model's minimum supported resolution is higher than 1024×1024, we use the smallest supported resolution closest to 1024×1024.
  • API mode: we use the provider's synchronous endpoint when one is available. For async-only providers, we poll the job status every 100 milliseconds so the measured queue wait reflects actual provider time, not poll granularity.
  • Aggregation: median, p05, p25, p75, and p95 are computed from the raw distribution of successful measurements with no outlier trimming.
  • Disabled provider features: watermarks and safety checks are disabled where the API supports it, to remove sources of latency variance unrelated to generation speed.
  • What is not measured: provider cold-start delay, authentication handshakes, retry overhead, and any client warm-up. Each entry in the dataset is a single one-shot measurement.
  • Infrastructure: runs on Google Cloud in us-central1.
  • New endpoints: for newly added endpoints, including those benchmarked ahead of public availability, we may adjust benchmarking cadence to build a distribution of generation times before results are listed. Published results are calculated over the trailing 3 days of measurements, so figures for newly listed endpoints reflect initial runs and may change as further measurements accumulate.

Model & Provider Inclusion Criteria

Our objective is to analyze and compare popular and high-performing image generation models to help users choose between them. As such, we apply tests for industry significance and competitive performance to evaluate the inclusion of new models and providers. We continuously refine these criteria and welcome feedback or suggestions. To suggest models or providers, please contact us via the contact page.

Statement of Independence

Benchmarking is conducted with strict independence and objectivity. No compensation is received from any providers for listing or for favorable outcomes on Artificial Analysis.