Best AI for video generation
Human preference over generated clips, from a prompt, from a still, and from an edit.
Rumeqo runs these models inside your team rooms. See what each one costs.
| rank | model | vendor | composite | benchmarks | elo | elo | elo |
|---|---|---|---|---|---|---|---|
| 1 | Dreamina Seedance 2.0 720pbytedance/dreamina-seedance-2.0-720p | ByteDance | 78.4 | 3 of 3 benchmarks | 1478 | 1478 | 1377 |
| 2 | Gemini Omni Flashgoogle/gemini-omni-flash | 74.6 | 3 of 3 benchmarks | 1512 | 1462 | 1347 | |
| 3 | Minimax H3minimax/minimax-h3 | MiniMax | 71.2 | 2 of 3 benchmarks | 1453 | 1476 | |
| 4 | HappyHorse 1.0alibaba/happyhorse-1.0 | Alibaba | 68.3 | 3 of 3 benchmarks | 1428 | 1442 | 1308 |
| 5 | Veo 3.1 Audiogoogle/veo-3.1-audio | 65.8 | 2 of 3 benchmarks | 1364 | 1397 | ||
| 6 | FLUX.3 Videoblack-forest-labs/flux-3-video | Black Forest Labs | 64.8 | 1 of 3 benchmarks | 1496 | ||
| 7 | Veo 3.1 Audio 1080pgoogle/veo-3.1-audio-1080p | 64.5 | 2 of 3 benchmarks | 1363 | 1390 | ||
| 8 | Grok Imagine Video 1.5 720px-ai/grok-imagine-video-1.5-720p | xAI | 63.9 | 1 of 3 benchmarks | 1462 | ||
| 9 | Grok Imagine Video 720px-ai/grok-imagine-video-720p | xAI | 63.4 | 2 of 3 benchmarks | 1347 | 1417 | |
| 10 | Veo 3.1 Fast Audiogoogle/veo-3.1-fast-audio | 63.3 | 2 of 3 benchmarks | 1362 | 1384 | ||
| 11 | Muse Videometa/muse-video | Meta | 63.2 | 1 of 3 benchmarks | 1457 | ||
| 12 | Veo 3.1 Fast Audio 1080pgoogle/veo-3.1-fast-audio-1080p | 61.5 | 2 of 3 benchmarks | 1358 | 1371 | ||
| 13 | Wan2.7 I2Valibaba/wan2.7-i2v | Alibaba | 61.5 | 1 of 3 benchmarks | 1427 | ||
| 14 | Sora 2 Proopenai/sora-2-pro | OpenAI | 60.8 | 1 of 3 benchmarks | 1366 | ||
| 15 | Veo 3 Fast Audiogoogle/veo-3-fast-audio | 58.5 | 2 of 3 benchmarks | 1347 | 1324 | ||
| 16 | Grok Imagine Video 480px-ai/grok-imagine-video-480p | xAI | 57.5 | 1 of 3 benchmarks | 1383 | ||
| 17 | Veo 3 Audiogoogle/veo-3-audio | 56.7 | 2 of 3 benchmarks | 1339 | 1330 | ||
| 18 | Vidu Q3 Proshengshu/vidu-q3-pro | Shengshu | 55.8 | 1 of 3 benchmarks | 1361 | ||
| 19 | Wan2.7 T2Valibaba/wan2.7-t2v | Alibaba | 55.3 | 1 of 3 benchmarks | 1343 | ||
| 20 | Kling v3 Prokwaivgi/kling-v3-pro | Kling AI | 55.0 | 1 of 3 benchmarks | 1359 | ||
| 21 | Sora 2openai/sora-2 | OpenAI | 54.5 | 1 of 3 benchmarks | 1340 | ||
| 22 | Seedance v1.5 Probytedance/seedance-v1.5-pro | ByteDance | 53.1 | 2 of 3 benchmarks | 1256 | 1307 | |
| 23 | Wan2.6 T2Valibaba/wan2.6-t2v | Alibaba | 52.9 | 1 of 3 benchmarks | 1333 | ||
| 24 | Wan2.5 I2V Previewalibaba/wan2.5-i2v-preview | Alibaba | 52.6 | 1 of 3 benchmarks | 1322 | ||
| 25 | Wan2.6 I2Valibaba/wan2.6-i2v | Alibaba | 51.8 | 1 of 3 benchmarks | 1311 | ||
| 26 | Wan2.5 T2V Previewalibaba/wan2.5-t2v-preview | Alibaba | 50.5 | 1 of 3 benchmarks | 1249 | ||
| 27 | Pixverse v5.6pixverse/pixverse-v5.6 | PixVerse | 50.1 | 2 of 3 benchmarks | 1240 | 1299 | |
| 28 | Grok Imagine Videox-ai/grok-imagine-video | SpaceXAI | 48.9 | 1 of 3 benchmarks | 1264 | ||
| 29 | Veo 3google/veo-3 | 48.2 | 2 of 3 benchmarks | 1253 | 1256 | ||
| 30 | Gen-4.5runway/gen-4.5 | Runway | 48.1 | 1 of 3 benchmarks | 1223 | ||
| 31 | Kling 2.6 Prokwaivgi/kling-2.6-pro | Kling AI | 47.7 | 2 of 3 benchmarks | 1217 | 1293 | |
| 32 | Kling 2.5 Turbo 1080pkwaivgi/kling-2.5-turbo-1080p | Kling AI | 47.7 | 2 of 3 benchmarks | 1219 | 1274 | |
| 33 | Veo 3 Fastgoogle/veo-3-fast | 47.6 | 2 of 3 benchmarks | 1248 | 1256 | ||
| 34 | Vidu Q2 Turboshengshu/vidu-q2-turbo | Shengshu | 44.4 | 1 of 3 benchmarks | 1242 | ||
| 35 | Hailuo 2.3minimax/hailuo-2.3 | MiniMax | 44.1 | 2 of 3 benchmarks | 1203 | 1260 | |
| 36 | Seedance v1 Probytedance/seedance-v1-pro | ByteDance | 43.5 | 2 of 3 benchmarks | 1190 | 1272 | |
| 37 | Kling O3 Prokwaivgi/kling-o3-pro | Kling AI | 43.4 | 1 of 3 benchmarks | 1251 | ||
| 38 | Kling v2.1 Standardkwaivgi/kling-v2.1-standard | Kling AI | 42.0 | 1 of 3 benchmarks | 1227 | ||
| 39 | Kandinsky 5.0 T2V Prosber/kandinsky-5.0-t2v-pro | Sber AI | 41.0 | 1 of 3 benchmarks | 1172 | ||
| 40 | Hailuo 02 Prominimax/hailuo-02-pro | MiniMax | 40.4 | 2 of 3 benchmarks | 1198 | 1227 | |
| 41 | Ray 3luma/ray-3 | Luma AI | 40.4 | 2 of 3 benchmarks | 1205 | 1225 | |
| 42 | Vidu Q2 Proshengshu/vidu-q2-pro | Shengshu | 39.6 | 1 of 3 benchmarks | 1222 | ||
| 43 | Kling O1 Prokwaivgi/kling-o1-pro | Kling AI | 38.5 | 2 of 3 benchmarks | 1205 | 1203 | |
| 44 | Hailuo 02 Fastminimax/hailuo-02-fast | MiniMax | 37.9 | 1 of 3 benchmarks | 1193 | ||
| 45 | Kling v2.1 Masterkwaivgi/kling-v2.1-master | Kling AI | 37.5 | 2 of 3 benchmarks | 1162 | 1234 | |
| 46 | Hailuo 02 Standardminimax/hailuo-02-standard | MiniMax | 37.4 | 2 of 3 benchmarks | 1180 | 1222 | |
| 47 | Kandinsky 5.0 T2V Litesber/kandinsky-5.0-t2v-lite | Sber AI | 36.2 | 1 of 3 benchmarks | 1113 | ||
| 48 | Hunyuan Video 1.5tencent/hunyuan-video-1.5 | Tencent | 35.0 | 2 of 3 benchmarks | 1169 | 1196 | |
| 49 | Soraopenai/sora | OpenAI | 34.6 | 1 of 3 benchmarks | 1069 | ||
| 50 | Runway Gen4 Turborunway/runway-gen4-turbo | Runway | 33.1 | 1 of 3 benchmarks | 1051 | ||
| 51 | Mochi v1genmo/mochi-v1 | Genmo | 32.2 | 1 of 3 benchmarks | 1005 | ||
| 52 | Runway Gen4 Alephrunway/runway-gen4-aleph | Runway | 32.2 | 1 of 3 benchmarks | 1194 | ||
| 53 | Veo 2google/veo-2 | 32.0 | 2 of 3 benchmarks | 1164 | 1164 | ||
| 54 | Wan v2.2 A14balibaba/wan-v2.2-a14b | Alibaba | 30.8 | 2 of 3 benchmarks | 1131 | 1169 | |
| 55 | Seedance v1 Litebytedance/seedance-v1-lite | ByteDance | 30.2 | 2 of 3 benchmarks | 1112 | 1184 | |
| 56 | Ltx 2 19Blightricks/ltx-2-19b | Lightricks | 30.2 | 2 of 3 benchmarks | 1151 | 1151 | |
| 57 | Ray2luma/ray2 | Luma AI | 26.6 | 2 of 3 benchmarks | 1064 | 1106 | |
| 58 | Pika v2.2pika/pika-v2.2 | Pika | 24.8 | 2 of 3 benchmarks | 1008 | 996 |
How this ranks
Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.
A model scored on fewer than 2 of the 3 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.
Data sources
Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.
- LMArenaCC BY 4.0
Arena ratings by LMArena, from the public leaderboard dataset.