Image Generation Benchmarking Methodology
Scope & Background
Artificial Analysis performs benchmarking on image generation models typically delivered via serverless API endpoints. This page describes the methodology used to measure the quality, speed, and price of image generation models.
For image generation capabilities, we cover two modalities:
- Text to Image: models that generate an image from a text prompt only.
- Image Editing: models that modify a reference image based on a text prompt.
Endpoint Coverage
Covered workloads: inference benchmarking covers Text to Image generation endpoints.
Public serverless endpoints only: benchmarking covers true serverless endpoints that are, or will be, publicly available. Demo products, hardware showcases, and dedicated or private deployments are out of scope.
Standard & Modified Endpoints
We categorize each endpoint as either Standard — serving the native model at full fidelity, with its standard input parameters exposed — or Modified — altered from the native model in ways that affect output fidelity, typically to increase generation speed or reduce cost.
Modified endpoints are labeled as Modified on Artificial Analysis and may be hidden from the default leaderboard view, remaining visible to users who opt to display them. To preserve the fairness of our default benchmarks, we retain discretion over whether an endpoint is categorized as Modified — indicators include distillation, partial exposure of the native model's input parameters, or materially degraded output quality — and over how many variants of a single model are listed.
Default Generation Settings
We generate images using each model's published defaults. For first-party APIs, this means the API's documented defaults; for open-source models, the defaults from the open-source repository. This covers inference steps, guidance scale, and any default negative prompt behavior. To enable fair comparison across models, we apply the following normalizations:
- We generate at the highest resolution each model supports, then downscale to 1024×1024 to serve in our Arenas.
- We generate 1 image per prompt.
- We use a 1:1 aspect ratio.
- We use a seed of 42.
For Image Editing models, each prompt is paired with a reference image from a curated set.
Key Metrics
We use the following metrics to track quality, performance and price for image generation models.
Quality Elo
Each model's Elo score reflects its relative quality, derived from user votes in the Image Arena. We compute ratings using Bradley-Terry Maximum Likelihood Estimation and rescale them to an Elo-like range for readability.
Each modality is scored independently, as Arena matchups only pair outputs from the same modality.
Price per 1,000 Images
Provider's price (USD) per image generated, multiplied by 1,000.
Pricing is taken directly from each provider's published per-image rate.
Generation Time
Median time the provider takes to generate a single image, calculated over the past 14 days.
Images are generated at a batch size of 1. Generation Time includes downloading the image from the provider where a URL is provided rather than an image response. This reflects end-user latency, since URLs can be issued before generation completes.
Text to Image
Methodology Overview
Text-to-image models vary in capability across a wide range of use cases. One model may lead on text rendering, another on complex layouts or photorealistic human characters. The model an agency chooses for campaign creative is not necessarily the one a product manager chooses for UI mockups, or a game studio for concept art.
Our methodology measures this variation directly. We rank models on human preference across a structured taxonomy of real-world use cases and model capabilities, identifying the best model overall and the best model for specific work. The taxonomy anticipates where the field is going, covering the capabilities leading labs are actively pursuing and the use cases where image generation is seeing real adoption. Our prompt set is refreshed monthly so the leaderboard continues to measure the frontier as it moves.
Prompt Curation and Refresh
We maintain a rotation of prompts that are refreshed regularly to stay at the frontier of rapidly evolving model capabilities, and to ensure fairness and accuracy of our leaderboard.
- Prompt Taxonomy: Every prompt is authored against a taxonomy of two axes, real-world use case and model capability, and tagged for both. Prompts are sampled evenly across the taxonomy in our arena.
- Real User Data: Our prompts are written to reflect how end users actually prompt, informed by anonymized crowdsourced data and a human-curated, continuously updated prompt corpus.
- Freshness: We refresh our prompt set every month. Prompts are retired based on how well they discriminate between models and how well they reflect the way real users prompt image models today.
Our prompt set carries the signal of real user data while providing unbiased, even coverage across use cases and capabilities.
Prompt Taxonomy
We construct our prompt taxonomy along two axes.
Use Case. These use cases are grounded in our analysis of how consumers and enterprises are adopting text-to-image generation, and our observation of emerging use cases. Example tasks for each use case include:
- Marketing & Advertising: Ad creatives, campaign visuals, posters, brand imagery
- Retail & E-commerce: Product shots, on-model apparel, packaging mockups
- Live-Action Film: Cinematic scenes, film stills, storyboards, character sheets
- Animation & Gaming: Game assets, concept art, character design, comics
- Architecture & Real Estate: Interior and exterior visualization, floor plans, staging
- Productivity & Knowledge Work: Diagrams, infographics, charts, slides
- UI/UX Design: UI mockups across app, web, in-car, and spatial surfaces
- Consumer: Book covers, stock images, editorial illustration
- Social Media & Creator Content: Thumbnails, promo cards, channel and profile art
- Frontier: Capabilities at and beyond the current frontier
Capability. These model capabilities are informed by our ongoing work with leading model labs and the capabilities they are pursuing, and grounded in academic benchmarks. Examples for each capability include:
- Reasoning: Entity, mathematical, spatial, logical reasoning, concept mixing, idiom interpretation
- Knowledge: Real landmarks, species, and domain facts across science and common sense
- Text Rendering: Long text, small text, symbols, artistic lettering
- Layout: Flows, arrows, blocks, visual hierarchy, multi-panel compositions
- Complex Compositions: Precise counting, spatial relationships, subject interactions, attribute binding
- Lighting: Reflection, refraction, shadows, caustics
- Material: Surface properties, transparency, subsurface scattering, texture realism
- Physics: Gravity, support, collision, thermal and state change
- Human Anatomy: Hands, faces, body proportion, dynamic poses and movement
Prompt Writing
Prompts follow strict authoring standards: plain natural language, a single positive prompt with the negative-prompt field left empty, and architecture-neutral phrasing that does not favor any model family. Prompts are screened for redundancy at authoring time to make sure they cover a diverse range of real world usage scenarios. Prompts are written:
- For full coverage of the taxonomy, distributed evenly across every use case and capability combination. Because each prompt carries a use case, capability, and style tag, we can evaluate not just overall model quality but performance on the specific cuts that matter to each user. One user needs the best model for UI design with complex layout hierarchy, another might need stunning cinematic shots with realistic human characters.
- To reflect how end users really prompt, informed by crowdsourced anonymized prompts from real users and a human-curated, live-updated prompt corpus of consumer prompting patterns.
All prompts are:
- English language
- Human curated
This results in prompt sets that are high signal and low noise, and that evaluate models against the generation tasks most relevant to end users and industry, today and in the near future.
Prompt Retirement
Every month, we select prompts to retire based on the following two criteria:
- Discrimination: whether each prompt still produces clear signal about model capability differences. For each vote, we label the higher-rated model (per the global leaderboard) the favorite. Prompts whose favorite win rate is statistically indistinguishable from chance are flagged for retirement.
- Freshness/Realism: we compare our benchmark prompt set against our live prompt corpus, including our latest crowdsourced prompts and human-curated prompt corpus, to filter out prompts that no longer reflect real-world prompting conventions.
Retired prompts are no longer served in the arena for vote collection. A monthly refresh cadence also protects the integrity of the leaderboard: a prompt set that rotates cannot be overfit.
Sampling Methodology
Each prompt in our prompt set is tagged with 1 use case, and 1 primary capability tag. We sample prompts evenly across each use case and capability combination.
Every matchup pairs two outputs generated from the identical prompt, within the same modality. Left/right position is randomized and model identities are revealed only after the vote. Newly added models are temporarily upsampled until their ratings converge, and opponents are drawn to maximize the information value of each matchup. Each use case and capability contributes equal weight to the overall leaderboard rating.
Elo Calculation
Vote Collection
Model quality is measured by human preference.
- Blind pairwise voting: Evaluators see two outputs generated from the same prompt by two different models, without model identities, and select the one they prefer.
- Judging hints: Each matchup surfaces short hints tied to the prompt's use case, capability, and style tags, directing attention to the specific aspects under test on complex prompts.
- Engagement gate: Votes can only be cast after a minimum engagement time with each output.
- Vote quality: All votes pass through bot detection and anomaly filtering before entering the rating calculation.
Vote Filtering
We filter votes according to what our current methodology measures. Votes that predate the current methodology were cast on prompts written to separate an earlier generation of models. Current models saturate many of those prompts, with nearly all producing an acceptable result, so a vote on them records a coin flip rather than a capability difference. We therefore apply a cohort-based filter to historical votes:
- Current cohort: Models that are publicly accessible, recently released, and of most relevance to our audience. Ranked on votes collected under the current methodology only.
- Legacy cohort: All other models. Ranked on every vote they appear in, preserving their representation on the leaderboard.
- Rule: A vote is retained if and only if both participating models are eligible for it.
This filter is applied after our standard vote-quality filters, including bot and spam exclusion. Elo is recalculated from scratch over the entire retained vote set, with every model, current cohort and legacy alike, ranked in that single calculation. The rating is anchored on FLUX.1 [schnell] = 1000 for the overall leaderboard, and FLUX.2 [dev] = 1000 for all subcategory leaderboards.
Generation Time Testing Methodology
Key technical details:
- Cadence: benchmarks run 4 times per day at random times.
- Sample size: the headline Generation Time is the median of successful measurements from the trailing 14 days, typically around 56 samples per host model per window.
- What is included: each measurement times the full request lifecycle for one image, covering the API request, provider inference, and either the inline response body or the URL download to disk. Where providers return a URL, the URL is often emitted before the image is fully written, so downloading the bytes is counted to reflect real end-user latency.
- Unique prompts: a new prompt is generated for each call from a curated pool, so providers cannot return cached responses across runs.
- Resolution: we generate at 1024×1024. If a model's minimum supported resolution is higher than 1024×1024, we use the smallest supported resolution closest to 1024×1024.
- API mode: we use the provider's synchronous endpoint when one is available. For async-only providers, we poll the job status every 100 milliseconds so the measured queue wait reflects actual provider time, not poll granularity.
- Aggregation: median, p05, p25, p75, and p95 are computed from the raw distribution of successful measurements with no outlier trimming.
- Disabled provider features: watermarks and safety checks are disabled where the API supports it, to remove sources of latency variance unrelated to generation speed.
- What is not measured: provider cold-start delay, authentication handshakes, retry overhead, and any client warm-up. Each entry in the dataset is a single one-shot measurement.
- Infrastructure: runs on Google Cloud in us-central1.
- New endpoints: for newly added endpoints, including those benchmarked ahead of public availability, we may adjust benchmarking cadence to build a distribution of generation times before results are listed. Published results are calculated over the trailing 14 days of measurements, so figures for newly listed endpoints reflect initial runs and may change as further measurements accumulate.
Model & Provider Inclusion Criteria
Our objective is to analyze and compare popular and high-performing image generation models to help users choose between them. As such, we apply tests for industry significance and competitive performance to evaluate the inclusion of new models and providers. We continuously refine these criteria and welcome feedback or suggestions. To suggest models or providers, please contact us via the contact page.
Statement of Independence
Benchmarking is conducted with strict independence and objectivity. No compensation is received from any providers for listing or for favorable outcomes on Artificial Analysis.