Image to Video Leaderboard (With Audio)Artificial Analysis

Added to the leaderboard in the last month:

Vidu Q3 Turbo, MiniMax H3, MAGI-2 Preview

Range
Creator
Model
Elo
95% CI
Samples
Released
API Pricing 1
11-2
ByteDance Seed logoByteDance Seed
Dreamina Seedance 2.0 720p
1,196-7/713,043Mar 2026$9.07 /min
21-3
MiniMax logoMiniMax
MiniMax H3Hugging FaceOpen Weights
1,190-9/95,915Jul 2026$7.80 /min
32-3
Google logoGoogle
Gemini Omni Flash
1,187-8/87,539May 2026$6.00 /min
44-6
SpaceXAI logoSpaceXAI
grok-imagine-video-1.5
1,111-8/85,276May 2026$8.40 /min
54-6
Alibaba-ATH logoAlibaba-ATH
HappyHorse-1.1
1,110-8/88,483Jun 2026$9.90 /min
64-6
Sand.ai logoSand.ai
MAGI-2 PreviewHugging FaceOpen Weights
1,105-8/810,910Aug 2026Coming soon
77-10
Alibaba logoAlibaba
Wan 2.7
1,090-8/84,577Apr 2026$9.00 /min
87-10
Alibaba-ATH logoAlibaba-ATH
HappyHorse-1.0
1,090-8/87,878Apr 2026$13.20 /min
97-10
Skywork AI logoSkywork AI
SkyReels V4
1,087-8/84,835Mar 2026$21.00 /min
107-11
Google logoGoogle
Veo 3.1
1,083-7/77,992Jan 2026$24.00 /min
1110-13
SpaceXAI logoSpaceXAI
grok-imagine-video
1,079-7/712,942Jan 2026$4.20 /min
1211-14
KlingAI logoKlingAI
Kling 3.0 1080p (Pro)
1,075-7/712,520Feb 2026$20.16 /min
1311-15
Google logoGoogle
Veo 3.1 Fast
1,073-7/711,894Jan 2026$9.00 /min
1412-17
PixVerse logoPixVerse
PixVerse V6
1,069-7/79,274Mar 2026$6.90 /min
1513-18
KlingAI logoKlingAI
Kling 3.0 720p (Standard)
1,067-7/712,558Feb 2026$15.60 /min
1613-18
Google logoGoogle
Veo 3.1 Lite
1,066-8/87,259Mar 2026$4.80 /min
1714-18
Vidu logoVidu
Vidu Q3 Pro
1,063-7/712,876Jan 2026$9.60 /min
1815-18
KlingAI logoKlingAI
Kling 3.0 Omni 1080p (Pro)
1,061-7/77,117Feb 2026$16.80 /min
1919
KlingAI logoKlingAI
Kling 3.0 Omni 720p (Standard)
1,052-7/77,005Feb 2026$13.44 /min
2019-20
Vidu logoVidu
Vidu Q3 Turbo
1,044-10/102,459Feb 2026$3.90 /min
2121-22
KlingAI logoKlingAI
Kling 2.6 Pro (January)
1,005-8/87,494Jan 2026$8.40 /min
2222
ByteDance Seed logoByteDance Seed
Seedance 1.5 pro
1,0000/09,124Dec 2025$11.86 /min
2323-25
Lightricks logoLightricks
LTX-2.3 FastHugging FaceOpen Weights
958-7/78,969Mar 2026$2.40 /min
2423-25
PixVerse logoPixVerse
PixVerse V5.6
955-8/86,189Feb 2026Coming soon
2523-25
Lightricks logoLightricks
LTX-2.3 ProHugging FaceOpen Weights
954-7/79,080Mar 2026$4.80 /min
2626-27
Lightricks logoLightricks
LTX-2 FastHugging FaceOpen Weights
933-8/85,928Oct 2025$2.40 /min
2726-27
Sapiens AI logoSapiens AI
Agnes-Video-V2.0
925-9/95,378May 2026$0.30 /min
2828
Alibaba logoAlibaba
Wan 2.6
892-8/86,112Dec 2025$9.00 /min
2929
Lightricks logoLightricks
LTX-2 ProHugging FaceOpen Weights
877-9/95,840Oct 2025$3.60 /min

1 API Pricing reflects the cost to generate 1 minute of 1080p video on the model creator's API at the model's default settings

Frequently Asked Questions

Dreamina Seedance 2.0 720p currently leads among Image to Video models with audio output in the Artificial Analysis Image to Video Arena with an Elo score of 1196.

The top Image to Video models with audio by Elo rating are: 1. Dreamina Seedance 2.0 720p (Elo 1196), 2. MiniMax H3 (Elo 1190), 3. Gemini Omni Flash (Elo 1187), 4. grok-imagine-video-1.5 (Elo 1111), 5. HappyHorse-1.1 (Elo 1110). Rankings are based on blind user votes in the Artificial Analysis Video Arena.

MiniMax H3 currently leads among open weights Image to Video models with audio in the Artificial Analysis Image to Video Arena with an Elo score of 1190, followed by MAGI-2 Preview (Elo 1105) and LTX-2.3 Fast (Elo 958).

Gemini Omni Flash currently leads the Artificial Analysis Image to Video Arena (without audio) with an Elo score of 1366.

The top Image to Video models without audio by Elo rating are: 1. Gemini Omni Flash (Elo 1366), 2. MiniMax H3 (Elo 1352), 3. Dreamina Seedance 2.0 720p (Elo 1337), 4. grok-imagine-video-1.5 (Elo 1329), 5. PixVerse V6 (Elo 1327). Rankings are based on blind user votes in the Artificial Analysis Video Arena.

MiniMax H3 currently leads among open weights Image to Video models without audio in the Artificial Analysis Image to Video Arena with an Elo score of 1352, followed by Cosmos3-Super-Image2Video-4Step (Elo 1264) and Cosmos3-Super-Image2Video (Elo 1246).

Text to Video models generate videos from text descriptions alone, while Image to Video models take an existing image as input and animate or extend it into a video. This allows for more control over the visual content and style of the generated video.

Models are ranked using an Elo rating system derived from user votes in blind comparisons. Users compare videos generated from the same input image and choose the result they prefer. Higher Elo scores indicate a model is preferred more often. Vote in the Artificial Analysis Video Arena