Image to Video Leaderboard (With Audio)Artificial Analysis

Added to the leaderboard in the last month:

Vidu Q3 Turbo, MiniMax H3, MAGI-2 Preview

Range
Creator
Model
Elo
95% CI
Samples
Released
API Pricing 1
11-2
ByteDance Seed logoByteDance Seed
Dreamina Seedance 2.0 720p
1,199-7/712,285Mar 2026$9.07 /min
21-3
MiniMax logoMiniMax
MiniMax H3Hugging FaceOpen Weights
1,194-10/105,649Jul 2026$7.80 /min
31-3
Google logoGoogle
Gemini Omni Flash
1,192-9/96,735May 2026$6.00 /min
44-6
SpaceXAI logoSpaceXAI
grok-imagine-video-1.5
1,116-8/84,978May 2026$8.40 /min
54-6
Alibaba-ATH logoAlibaba-ATH
HappyHorse-1.1
1,111-8/87,506Jun 2026$9.90 /min
64-6
Sand.ai logoSand.ai
MAGI-2 Preview
1,109-8/810,631Aug 2026Coming soon
77-10
Alibaba logoAlibaba
Wan 2.7
1,094-9/94,279Apr 2026$9.00 /min
87-10
Alibaba-ATH logoAlibaba-ATH
HappyHorse-1.0
1,091-8/87,619Apr 2026$13.20 /min
97-11
Skywork AI logoSkywork AI
SkyReels V4
1,088-8/84,565Mar 2026$21.00 /min
108-11
Google logoGoogle
Veo 3.1
1,086-7/77,731Jan 2026$24.00 /min
1110-13
SpaceXAI logoSpaceXAI
grok-imagine-video
1,081-7/711,997Jan 2026$4.20 /min
1211-14
Google logoGoogle
Veo 3.1 Fast
1,078-7/710,983Jan 2026$9.00 /min
1311-14
KlingAI logoKlingAI
Kling 3.0 1080p (Pro)
1,077-7/711,619Feb 2026$20.16 /min
1412-18
PixVerse logoPixVerse
PixVerse V6
1,071-7/79,035Mar 2026$6.90 /min
1514-18
KlingAI logoKlingAI
Kling 3.0 720p (Standard)
1,069-7/711,565Feb 2026$15.60 /min
1614-18
Google logoGoogle
Veo 3.1 Lite
1,066-8/87,022Mar 2026$4.80 /min
1714-18
KlingAI logoKlingAI
Kling 3.0 Omni 1080p (Pro)
1,065-7/76,895Feb 2026$16.80 /min
1814-18
Vidu logoVidu
Vidu Q3 Pro
1,065-7/711,967Jan 2026$9.60 /min
1919
KlingAI logoKlingAI
Kling 3.0 Omni 720p (Standard)
1,056-7/76,807Feb 2026$13.44 /min
2020
Vidu logoVidu
Vidu Q3 Turbo
1,043-10/102,195Feb 2026$3.90 /min
2121-22
KlingAI logoKlingAI
Kling 2.6 Pro (January)
1,002-8/87,307Jan 2026$8.40 /min
2222
ByteDance Seed logoByteDance Seed
Seedance 1.5 pro
1,0000/08,382Dec 2025$11.86 /min
2323-25
Lightricks logoLightricks
LTX-2.3 FastHugging FaceOpen Weights
960-8/88,388Mar 2026$2.40 /min
2423-25
Lightricks logoLightricks
LTX-2.3 ProHugging FaceOpen Weights
959-8/88,492Mar 2026$4.80 /min
2523-25
PixVerse logoPixVerse
PixVerse V5.6
956-8/85,999Feb 2026Coming soon
2626-27
Lightricks logoLightricks
LTX-2 FastHugging FaceOpen Weights
933-8/85,818Oct 2025$2.40 /min
2726-27
Sapiens AI logoSapiens AI
Agnes-Video-V2.0
930-9/95,145May 2026$0.30 /min
2828
Alibaba logoAlibaba
Wan 2.6
896-9/95,955Dec 2025$9.00 /min
2929
Lightricks logoLightricks
LTX-2 ProHugging FaceOpen Weights
878-9/95,769Oct 2025$3.60 /min

1 API Pricing reflects the cost to generate 1 minute of 1080p video on the model creator's API at the model's default settings

Frequently Asked Questions

Dreamina Seedance 2.0 720p currently leads among Image to Video models with audio output in the Artificial Analysis Image to Video Arena with an Elo score of 1199.

The top Image to Video models with audio by Elo rating are: 1. Dreamina Seedance 2.0 720p (Elo 1199), 2. MiniMax H3 (Elo 1194), 3. Gemini Omni Flash (Elo 1192), 4. grok-imagine-video-1.5 (Elo 1116), 5. HappyHorse-1.1 (Elo 1111). Rankings are based on blind user votes in the Artificial Analysis Video Arena.

MiniMax H3 currently leads among open weights Image to Video models with audio in the Artificial Analysis Image to Video Arena with an Elo score of 1194, followed by LTX-2.3 Fast (Elo 960) and LTX-2.3 Pro (Elo 959).

Gemini Omni Flash currently leads the Artificial Analysis Image to Video Arena (without audio) with an Elo score of 1369.

The top Image to Video models without audio by Elo rating are: 1. Gemini Omni Flash (Elo 1369), 2. MiniMax H3 (Elo 1351), 3. Dreamina Seedance 2.0 720p (Elo 1338), 4. grok-imagine-video-1.5 (Elo 1330), 5. grok-imagine-video (Elo 1326). Rankings are based on blind user votes in the Artificial Analysis Video Arena.

MiniMax H3 currently leads among open weights Image to Video models without audio in the Artificial Analysis Image to Video Arena with an Elo score of 1351, followed by Cosmos3-Super-Image2Video-4Step (Elo 1266) and Cosmos3-Super-Image2Video (Elo 1244).

Text to Video models generate videos from text descriptions alone, while Image to Video models take an existing image as input and animate or extend it into a video. This allows for more control over the visual content and style of the generated video.

Models are ranked using an Elo rating system derived from user votes in blind comparisons. Users compare videos generated from the same input image and choose the result they prefer. Higher Elo scores indicate a model is preferred more often. Vote in the Artificial Analysis Video Arena