Image to Video Leaderboard (With Audio)Artificial Analysis
Vidu Q3 Turbo, MiniMax H3, MAGI-2 Preview
Range | Creator | Model | Elo | 95% CI | Samples | Released | API Pricing 1 | |
|---|---|---|---|---|---|---|---|---|
| 1 | 1-2 | Dreamina Seedance 2.0 720p | 1,199 | -7/7 | 12,285 | Mar 2026 | $9.07 /min | |
| 2 | 1-3 | MiniMax H3 | 1,194 | -10/10 | 5,649 | Jul 2026 | $7.80 /min | |
| 3 | 1-3 | Gemini Omni Flash | 1,192 | -9/9 | 6,735 | May 2026 | $6.00 /min | |
| 4 | 4-6 | grok-imagine-video-1.5 | 1,116 | -8/8 | 4,978 | May 2026 | $8.40 /min | |
| 5 | 4-6 | HappyHorse-1.1 | 1,111 | -8/8 | 7,506 | Jun 2026 | $9.90 /min | |
| 6 | 4-6 | MAGI-2 Preview | 1,109 | -8/8 | 10,631 | Aug 2026 | Coming soon | |
| 7 | 7-10 | Wan 2.7 | 1,094 | -9/9 | 4,279 | Apr 2026 | $9.00 /min | |
| 8 | 7-10 | HappyHorse-1.0 | 1,091 | -8/8 | 7,619 | Apr 2026 | $13.20 /min | |
| 9 | 7-11 | SkyReels V4 | 1,088 | -8/8 | 4,565 | Mar 2026 | $21.00 /min | |
| 10 | 8-11 | Veo 3.1 | 1,086 | -7/7 | 7,731 | Jan 2026 | $24.00 /min | |
| 11 | 10-13 | grok-imagine-video | 1,081 | -7/7 | 11,997 | Jan 2026 | $4.20 /min | |
| 12 | 11-14 | Veo 3.1 Fast | 1,078 | -7/7 | 10,983 | Jan 2026 | $9.00 /min | |
| 13 | 11-14 | Kling 3.0 1080p (Pro) | 1,077 | -7/7 | 11,619 | Feb 2026 | $20.16 /min | |
| 14 | 12-18 | PixVerse V6 | 1,071 | -7/7 | 9,035 | Mar 2026 | $6.90 /min | |
| 15 | 14-18 | Kling 3.0 720p (Standard) | 1,069 | -7/7 | 11,565 | Feb 2026 | $15.60 /min | |
| 16 | 14-18 | Veo 3.1 Lite | 1,066 | -8/8 | 7,022 | Mar 2026 | $4.80 /min | |
| 17 | 14-18 | Kling 3.0 Omni 1080p (Pro) | 1,065 | -7/7 | 6,895 | Feb 2026 | $16.80 /min | |
| 18 | 14-18 | Vidu Q3 Pro | 1,065 | -7/7 | 11,967 | Jan 2026 | $9.60 /min | |
| 19 | 19 | Kling 3.0 Omni 720p (Standard) | 1,056 | -7/7 | 6,807 | Feb 2026 | $13.44 /min | |
| 20 | 20 | Vidu Q3 Turbo | 1,043 | -10/10 | 2,195 | Feb 2026 | $3.90 /min | |
| 21 | 21-22 | Kling 2.6 Pro (January) | 1,002 | -8/8 | 7,307 | Jan 2026 | $8.40 /min | |
| 22 | 22 | Seedance 1.5 pro | 1,000 | 0/0 | 8,382 | Dec 2025 | $11.86 /min | |
| 23 | 23-25 | LTX-2.3 Fast | 960 | -8/8 | 8,388 | Mar 2026 | $2.40 /min | |
| 24 | 23-25 | LTX-2.3 Pro | 959 | -8/8 | 8,492 | Mar 2026 | $4.80 /min | |
| 25 | 23-25 | PixVerse V5.6 | 956 | -8/8 | 5,999 | Feb 2026 | Coming soon | |
| 26 | 26-27 | LTX-2 Fast | 933 | -8/8 | 5,818 | Oct 2025 | $2.40 /min | |
| 27 | 26-27 | Agnes-Video-V2.0 | 930 | -9/9 | 5,145 | May 2026 | $0.30 /min | |
| 28 | 28 | Wan 2.6 | 896 | -9/9 | 5,955 | Dec 2025 | $9.00 /min | |
| 29 | 29 | LTX-2 Pro | 878 | -9/9 | 5,769 | Oct 2025 | $3.60 /min |
1 API Pricing reflects the cost to generate 1 minute of 1080p video on the model creator's API at the model's default settings
Frequently Asked Questions
Dreamina Seedance 2.0 720p currently leads among Image to Video models with audio output in the Artificial Analysis Image to Video Arena with an Elo score of 1199.
The top Image to Video models with audio by Elo rating are: 1. Dreamina Seedance 2.0 720p (Elo 1199), 2. MiniMax H3 (Elo 1194), 3. Gemini Omni Flash (Elo 1192), 4. grok-imagine-video-1.5 (Elo 1116), 5. HappyHorse-1.1 (Elo 1111). Rankings are based on blind user votes in the Artificial Analysis Video Arena.
MiniMax H3 currently leads among open weights Image to Video models with audio in the Artificial Analysis Image to Video Arena with an Elo score of 1194, followed by LTX-2.3 Fast (Elo 960) and LTX-2.3 Pro (Elo 959).
Gemini Omni Flash currently leads the Artificial Analysis Image to Video Arena (without audio) with an Elo score of 1369.
The top Image to Video models without audio by Elo rating are: 1. Gemini Omni Flash (Elo 1369), 2. MiniMax H3 (Elo 1351), 3. Dreamina Seedance 2.0 720p (Elo 1338), 4. grok-imagine-video-1.5 (Elo 1330), 5. grok-imagine-video (Elo 1326). Rankings are based on blind user votes in the Artificial Analysis Video Arena.
MiniMax H3 currently leads among open weights Image to Video models without audio in the Artificial Analysis Image to Video Arena with an Elo score of 1351, followed by Cosmos3-Super-Image2Video-4Step (Elo 1266) and Cosmos3-Super-Image2Video (Elo 1244).
Text to Video models generate videos from text descriptions alone, while Image to Video models take an existing image as input and animate or extend it into a video. This allows for more control over the visual content and style of the generated video.
Models are ranked using an Elo rating system derived from user votes in blind comparisons. Users compare videos generated from the same input image and choose the result they prefer. Higher Elo scores indicate a model is preferred more often. Vote in the Artificial Analysis Video Arena