Video Generation Benchmarking Methodology

Scope & Background

Artificial Analysis performs benchmarking on video generation models typically delivered via serverless API endpoints. This page describes the methodology used to measure the quality, speed, and price of video generation models.

For video generation capabilities, we publish six benchmarks:

  • AA-Video-T2V-Silent v2.0: models that generate video from a text prompt only.
  • AA-Video-I2V-Silent v1.0: models that generate video using a reference image as the first frame, and where supported an accompanying text prompt.
  • AA-Video-Editing-Silent v1.0: models that generate video from a video input and a text prompt.
  • AA-Video-T2V v2.0: Text to Video models that generate synchronized audio alongside the video.
  • AA-Video-I2V v1.0: Image to Video models that generate synchronized audio alongside the video.
  • AA-Video-Editing v1.0: Video Editing models that generate synchronized audio alongside the video.

Each benchmark has its own Elo pool in the Video Arena, so silent and audio-enabled outputs are not compared directly, and editing is not compared with generation.

We apply a standard set of default settings to every video we generate across all models and providers. We use videos generated with identical settings for each model, including resolution, frame rate, duration, aspect ratio, seed, guidance, and inference steps, to ensure comparability. The full default settings are listed in Default Generation Settings below.

Endpoint Coverage

Covered workloads: inference benchmarking covers Text to Video and Image to Video generation endpoints.

Public serverless endpoints only: benchmarking covers true serverless endpoints that are, or will be, publicly available. Demo products, hardware showcases, and dedicated or private deployments are out of scope.

Standard & Modified Endpoints

We categorize each endpoint as either Standard — serving the native model at full fidelity, with its standard input parameters exposed — or Modified — altered from the native model in ways that affect output fidelity, typically to increase generation speed or reduce cost.

Modified endpoints are labeled as Modified on Artificial Analysis and may be hidden from the default leaderboard view, remaining visible to users who opt to display them. To preserve the fairness of our default benchmarks, we retain discretion over whether an endpoint is categorized as Modified — indicators include distillation, partial exposure of the native model's input parameters, or materially degraded output quality — and over how many variants of a single model are listed.

Key Metrics

We use the following metrics to track quality, performance and price for video generation models.

Quality Elo

Each model's Elo score reflects its relative quality, derived from blind votes cast by our recruited panel, together with public votes cast in the Artificial Analysis Video Arena before 1 January 2026. In both cases, voters compare two videos generated from the same prompt by different models and select the one they prefer. Public votes from before that date exist only on Text to Video and Image to Video, and now count only for Image to Video: the Text to Video benchmarks were reset for v2.0 (see AA-Video-T2V v2.0 & AA-Video-T2V-Silent v2.0 below). Video Editing and the three modalities with audio collected their first public vote after it, so their scores rest on panel votes alone. We aggregate these pairwise votes using Bradley-Terry Maximum Likelihood Estimation and rescale them to an Elo-like range for readability. Ratings are recomputed hourly.

Elo is reported separately for each benchmark (AA-Video-T2V-Silent v2.0, AA-Video-I2V-Silent v1.0, AA-Video-Editing-Silent v1.0, AA-Video-T2V v2.0, AA-Video-I2V v1.0, AA-Video-Editing v1.0), as Arena matchups only pair outputs from the same modality.

Leaderboards

We publish one global leaderboard for each benchmark. AA-Video-T2V v2.0 and AA-Video-T2V-Silent v2.0 also publish category leaderboards (see below).

  • Global Leaderboard: the ranking published on this site, which aggregates every vote retained by the rules above. A model appears once it has received a minimum number of votes.

Price per Minute of Video

Provider's price (USD) to generate one minute of video at the default generation settings.

Pricing is taken directly from each provider's published per-minute rate.

Generation Time

Median time the provider takes to generate a single video at the default generation settings, calculated over the past 3 days.

Generation Time is measured end-to-end: from the moment the request is submitted to the provider, through any queue time, through inference and video encoding, through downloading the completed video from the provider. For asynchronous APIs (submit → poll → download), this includes all three phases.

For Image to Video models, Generation Time also includes any time required to upload the reference image to the provider. For providers that accept a reference-image URL, there is no upload step and Generation Time excludes image upload.

Default Generation Settings

These settings apply to both Text to Video and Image to Video benchmarks unless otherwise specified.

SettingValue
Resolution1080p (or closest supported value)
Frame rate24 FPS (or closest supported value)
Duration10 seconds (or equivalent number of frames when duration must be specified in frames)
Aspect ratio16:9 horizontal
Seed42 (where the model exposes a seed parameter)
Prompt enhancementOn when recommended by the model creator, otherwise off
CFG / guidanceFirst-party API default, or the open-source repository default
Inference stepsFirst-party API default, or the open-source repository default

Where a provider does not support the exact default, we use the closest supported value, breaking ties upward (e.g., when 24 FPS is not available, we prefer 30 FPS over 18 FPS since both are equidistant from the default). AA-Video-T2V v2.0 and AA-Video-I2V v1.0 benchmark models with audio output enabled; all other defaults apply.

Text to Video

Text to Video models take a text prompt as their only input. Prompts are drawn from a curated and reproducible prompt set.

Image to Video

Image to Video models take a reference image and, where supported, an accompanying text prompt. Reference images used in our benchmarks follow these standards:

  • Format: PNG (lossless). We do not use JPG/JPEG to avoid the compression artifacts lossy encoding can introduce.
  • Resolution: 1920×1080 (16:9, matching the target video resolution).
  • Byte fallback: Where a provider does not accept URLs, or requires a specific size or format, we download the image, apply the required transformations, and upload the bytes directly.
  • Upload time: Where a provider requires us to upload the image to their platform as a pre-generation step, that upload time is counted toward Generation Time.

AA-Video-T2V v2.0 & AA-Video-T2V-Silent v2.0

Methodology Overview

Text to video models vary in capability across use cases. One model may lead on human performance and lip sync, another on animation, on-screen text, or UI motion. The model a studio chooses for a brand film is not necessarily the one a product team chooses for UI motion demos, or a creator for social clips.

Our methodology measures this variation directly. We rank models on human preference across a taxonomy of real-world use cases and model capabilities, identifying the best model overall and the best model for specific work. We refresh our prompt set regularly so the leaderboard continues to measure the frontier as it moves.

AA-Video-T2V v2.0 judges text to video with audio. AA-Video-T2V-Silent v2.0 judges text to video without audio.

Prompt Curation and Refresh

We maintain a curated prompt set, refreshed regularly to track the frontier as model capabilities evolve.

  • Prompt Taxonomy: Every prompt is authored against two axes, real-world use case and model capability, and tagged for both. Each prompt is also tagged with its visual style. Prompts are sampled evenly across use case and capability in our arena.
  • Prompt Sets: AA-Video-T2V v2.0 uses 1,000 prompts. AA-Video-T2V-Silent v2.0 uses 500 prompts.
  • Freshness: We refresh our prompt set regularly. Prompts are retired when they no longer separate models or no longer reflect how people prompt video models today.

Prompt Taxonomy

We construct our prompt taxonomy along two axes, and tag visual style separately.

Use Case. These use cases are grounded in our analysis of how consumers and enterprises are adopting text to video generation, and our observation of emerging use cases. Example tasks for each include:

  • Marketing & Advertising: Animated posters, logo stings, brand films, hero product spots, launch videos
  • Retail & E-commerce: Studio product shots, product demos, apparel and fashion movement, creator-style testimonials, outfit try-ons
  • Live-Action Film: Cinematic shots, microdrama episodes, previsualization
  • Animation & Gaming: Animated scenes, game trailers, previsualization for animation and games
  • Architecture & Real Estate: Exterior flythroughs, interior walkthroughs, virtual staging
  • Productivity & Knowledge Work: Explainers, educational and training videos, animated data visualizations, scenario simulations
  • UI/UX & Motion Design: Kinetic typography, UI motion demos across devices, abstract motion graphics
  • Consumer: Selfies, personal moments, everyday stock footage and b-roll
  • Social Media & Creator Content: Talking heads, podcast clips, ASMR and satisfying videos, dance, vlogs, trending formats
  • Frontier: Capabilities at and beyond the current frontier

Capability. These model capabilities are informed by our ongoing work with leading model labs and the capabilities they are pursuing, and grounded in academic benchmarks. Examples for each include:

  • Physics: Mechanics, thermal effects, material interactions, fracture and deformation, optics, cause and effect, counterfactual physics
  • Complex Composition: Attribute binding, spatial relationships, multi-object interaction, counting, temporal ordering
  • Human Anatomy: Gait and locomotion, fine motor movement, facial expression, acrobatic motion, object grasping, human interaction
  • Lighting & Materials: Lighting, exposure, shadows, color temperature, texture, cloth, hair and fur
  • Multi-Scene & Narrative: Multi-event stories, shot transitions, montage, identity consistency across shots
  • Camera Control: Shot types such as close-up, aerial, POV and low angle, and moves such as tracking, dolly, orbit, crane, crash zoom, whip pan and rack focus
  • Text Rendering: Short and long text, 2D layout, non-lexical strings, text on deforming surfaces
  • Spatio-Temporal Consistency: Spatial coherence, time lapses, occlusion and re-entry, single-shot stories
  • Audio Synchronization: Diegetic sound effects, background music and score, on- and off-screen sound
  • Dialogue & Lip Sync: Single- and multi-speaker dialogue, voiceover vs on-screen speech

Style. Each prompt is also tagged with one of eight visual styles: Photorealistic, 3D Render/CGI, Cartoon and Anime, Flat Design, Hand-drawn/Illustration, Stop-Motion/Claymation, Lo-Fi, and Mixed Styles. Styles are not sampled evenly, and most prompts are photorealistic.

Prompt Writing

We write prompts to separate frontier video models, which handle simple requests well. Prompts are written:

  • For full coverage of the taxonomy, distributed evenly across every use case and capability combination, with a high complexity floor in each capability, such as large-scale rather than subtle motion and multiple interacting subjects.
  • To reflect real production work, grounded in model developers' prompting guides, our existing prompt sets, and the workflows we see across industry.
  • With sound in mind for AA-Video-T2V v2.0: prompts specify dialogue, sound effects, ambience, and music where the scene calls for them.

A human reviews every prompt before we serve it in the arena.

Prompt Retirement

We retire prompts when arena votes show they no longer separate models, or when they no longer reflect how people prompt video models.

Sampling Methodology

For each matchup, we sample evenly across use case and capability combinations, then draw a prompt at random within the chosen combination.

Every matchup pairs two videos generated from the same prompt, within the same benchmark, so videos with audio are never paired with silent ones. Left/right position is randomized.

Playback Standards

We serve every video to evaluators under the same standards:

  • Resolution: up to 1080p, as generated. Videos are never upscaled: models that cannot generate 1080p are served at their highest supported resolution, and output above 1080p is downscaled to 1080p.
  • Encoding: H.264 at high quality, with bitrate peaks capped at 10 Mbps.
  • Duration: 10 seconds at 16:9. Models that cannot generate 10 seconds are served at their maximum length.
  • Audio levelling: for AA-Video-T2V v2.0, we level each video's loudness toward -18 LUFS with a single fixed gain. A model's mix is never compressed or remastered, and no model wins by being louder.

Elo Calculation

Vote Collection

Model quality is measured by human preference.

  • Blind pairwise voting: Evaluators watch two videos generated from the same prompt by two different models, without model identities, and select the one they prefer. Both videos play at full size in a single player, and evaluators switch between them.
  • Judging hints: Each matchup surfaces short hints tied to the prompt's use case, primary capability, and style, directing attention to the specific aspects under test.
  • Engagement gate: Votes can only be cast after each video has played for at least 8 seconds. For AA-Video-T2V v2.0, evaluators are asked to judge the sound too, using headphones.
  • Vote quality: Both leaderboards are ranked on votes from our recruited panel only, screened and verified by the panel provider. About one in five matchups in each session is a quality check with a strong consensus answer. Sessions that fail these checks are rejected and all their votes removed, and quality-check votes are not counted in the rating.

Vote Filtering

AA-Video-T2V v2.0 and AA-Video-T2V-Silent v2.0 are new leaderboards. We retire every Text to Video vote cast before the refresh, for every model, because those votes were collected on earlier prompts and playback standards. Public Video Arena votes do not count toward these ratings. We calculate Elo from scratch over the retained votes, anchored on Kling 3.0 1080p (Pro) = 1000 for the overall leaderboard and for all subcategory leaderboards.

Leaderboards

For each benchmark we publish:

  • Overall Leaderboard: the ranking across every retained vote. A model appears once it has received a minimum number of votes.
  • Category Leaderboards: the same calculation restricted to prompts tagged with a given use case, capability, or style, published once the category has enough votes.

Generation Time Testing Methodology

Key technical details:

  • Cadence: Text to Video runs every 2 hours at a random offset within each window.
  • Sample size: the headline Generation Time is the median of successful measurements from the trailing 3 days. With measurements every 2 hours, Text to Video typically accumulates around 36 samples per host per window.
  • Prompt selection: each run draws a single random prompt from a curated database pool and submits it to every active host model.
  • Seed: a fixed seed of 42 is used where the model exposes a seed parameter, so outputs are reproducible across runs.
  • End-to-end timing: Generation Time is measured end-to-end from the first API call to the completed video download. For providers with asynchronous APIs (submit, poll, download), it includes queue wait, inference, and download.
  • API mode: we poll the job status every 100 milliseconds so the measured queue wait reflects actual provider time, not poll granularity.
  • Scope: Generation Time is currently measured for silent Text to Video and Image to Video models only. AA-Video-Editing-Silent v1.0 and the three audio benchmarks (AA-Video-T2V v2.0, AA-Video-I2V v1.0, AA-Video-Editing v1.0) are ranked via the Video Arena (Quality Elo) but are not currently latency-benchmarked.
  • Aggregation: median, p05, p25, p75, and p95 are computed from the raw distribution of successful measurements with no outlier trimming.
  • What is not measured: provider cold-start delay, authentication handshakes, retry overhead, and any client warm-up. Each entry in the dataset is a single one-shot measurement.
  • Infrastructure: runs on Google Cloud Run in us-central1.
  • New endpoints: for newly added endpoints, including those benchmarked ahead of public availability, we may adjust benchmarking cadence to build a distribution of generation times before results are listed. Published results are calculated over the trailing 3 days of measurements, so figures for newly listed endpoints reflect initial runs and may change as further measurements accumulate.

Model & Provider Inclusion Criteria

Our objective is to analyze and compare popular and high-performing video generation models to help users choose between them. As such, we apply tests for industry significance and competitive performance to evaluate the inclusion of new models and providers. We continuously refine these criteria and welcome feedback or suggestions. To suggest models or providers, please contact us via the contact page.

Statement of Independence

Benchmarking is conducted with strict independence and objectivity. No compensation is received from any providers for listing or for favorable outcomes on Artificial Analysis.