Image to Video Leaderboard (With Audio)Artificial Analysis

Added to the leaderboard in the last month:

Vidu Q3 Turbo, MiniMax H3, MAGI-2 Preview

Range
Creator
Model
Elo
95% CI
Samples
Released
API Pricing 1
11-2
ByteDance Seed logoByteDance Seed
Dreamina Seedance 2.0 720p
1,199-7/712,341Mar 2026$9.07 /min
21-3
MiniMax logoMiniMax
MiniMax H3Hugging FaceOpen Weights
1,193-10/105,679Jul 2026$7.80 /min
31-3
Google logoGoogle
Gemini Omni Flash
1,191-9/96,812May 2026$6.00 /min
44-6
SpaceXAI logoSpaceXAI
grok-imagine-video-1.5
1,115-8/85,011May 2026$8.40 /min
54-6
Alibaba-ATH logoAlibaba-ATH
HappyHorse-1.1
1,112-8/87,626Jun 2026$9.90 /min
64-6
Sand.ai logoSand.ai
MAGI-2 Preview
1,108-8/810,650Aug 2026Coming soon
77-10
Alibaba logoAlibaba
Wan 2.7
1,092-9/94,320Apr 2026$9.00 /min
87-10
Alibaba-ATH logoAlibaba-ATH
HappyHorse-1.0
1,090-8/87,653Apr 2026$13.20 /min
97-10
Skywork AI logoSkywork AI
SkyReels V4
1,088-8/84,603Mar 2026$21.00 /min
107-11
Google logoGoogle
Veo 3.1
1,086-7/77,758Jan 2026$24.00 /min
1110-13
SpaceXAI logoSpaceXAI
grok-imagine-video
1,079-7/712,153Jan 2026$4.20 /min
1211-13
Google logoGoogle
Veo 3.1 Fast
1,077-7/711,119Jan 2026$9.00 /min
1311-14
KlingAI logoKlingAI
Kling 3.0 1080p (Pro)
1,077-7/711,757Feb 2026$20.16 /min
1413-17
PixVerse logoPixVerse
PixVerse V6
1,070-7/79,066Mar 2026$6.90 /min
1514-18
KlingAI logoKlingAI
Kling 3.0 720p (Standard)
1,068-7/711,732Feb 2026$15.60 /min
1614-18
Google logoGoogle
Veo 3.1 Lite
1,066-8/87,053Mar 2026$4.80 /min
1714-18
Vidu logoVidu
Vidu Q3 Pro
1,064-7/712,124Jan 2026$9.60 /min
1815-18
KlingAI logoKlingAI
Kling 3.0 Omni 1080p (Pro)
1,063-7/76,927Feb 2026$16.80 /min
1919
KlingAI logoKlingAI
Kling 3.0 Omni 720p (Standard)
1,055-7/76,836Feb 2026$13.44 /min
2020
Vidu logoVidu
Vidu Q3 Turbo
1,042-10/102,247Feb 2026$3.90 /min
2121-22
KlingAI logoKlingAI
Kling 2.6 Pro (January)
1,003-8/87,363Jan 2026$8.40 /min
2222
ByteDance Seed logoByteDance Seed
Seedance 1.5 pro
1,0000/08,616Dec 2025$11.86 /min
2323-25
Lightricks logoLightricks
LTX-2.3 ProHugging FaceOpen Weights
959-8/88,704Mar 2026$4.80 /min
2423-25
Lightricks logoLightricks
LTX-2.3 FastHugging FaceOpen Weights
958-8/88,618Mar 2026$2.40 /min
2523-25
PixVerse logoPixVerse
PixVerse V5.6
957-8/86,087Feb 2026Coming soon
2626-27
Lightricks logoLightricks
LTX-2 FastHugging FaceOpen Weights
932-8/85,884Oct 2025$2.40 /min
2726-27
Sapiens AI logoSapiens AI
Agnes-Video-V2.0
930-9/95,254May 2026$0.30 /min
2828
Alibaba logoAlibaba
Wan 2.6
896-9/96,037Dec 2025$9.00 /min
2929
Lightricks logoLightricks
LTX-2 ProHugging FaceOpen Weights
880-9/95,818Oct 2025$3.60 /min

1 API Pricing reflects the cost to generate 1 minute of 1080p video on the model creator's API at the model's default settings

Frequently Asked Questions

Dreamina Seedance 2.0 720p currently leads among Image to Video models with audio output in the Artificial Analysis Image to Video Arena with an Elo score of 1199.

The top Image to Video models with audio by Elo rating are: 1. Dreamina Seedance 2.0 720p (Elo 1199), 2. MiniMax H3 (Elo 1193), 3. Gemini Omni Flash (Elo 1191), 4. grok-imagine-video-1.5 (Elo 1115), 5. HappyHorse-1.1 (Elo 1112). Rankings are based on blind user votes in the Artificial Analysis Video Arena.

MiniMax H3 currently leads among open weights Image to Video models with audio in the Artificial Analysis Image to Video Arena with an Elo score of 1193, followed by LTX-2.3 Pro (Elo 959) and LTX-2.3 Fast (Elo 958).

Gemini Omni Flash currently leads the Artificial Analysis Image to Video Arena (without audio) with an Elo score of 1369.

The top Image to Video models without audio by Elo rating are: 1. Gemini Omni Flash (Elo 1369), 2. MiniMax H3 (Elo 1351), 3. Dreamina Seedance 2.0 720p (Elo 1339), 4. grok-imagine-video-1.5 (Elo 1330), 5. grok-imagine-video (Elo 1326). Rankings are based on blind user votes in the Artificial Analysis Video Arena.

MiniMax H3 currently leads among open weights Image to Video models without audio in the Artificial Analysis Image to Video Arena with an Elo score of 1351, followed by Cosmos3-Super-Image2Video-4Step (Elo 1266) and Cosmos3-Super-Image2Video (Elo 1244).

Text to Video models generate videos from text descriptions alone, while Image to Video models take an existing image as input and animate or extend it into a video. This allows for more control over the visual content and style of the generated video.

Models are ranked using an Elo rating system derived from user votes in blind comparisons. Users compare videos generated from the same input image and choose the result they prefer. Higher Elo scores indicate a model is preferred more often. Vote in the Artificial Analysis Video Arena