llm leaderboard

Best AI for image generation

Human preference across generation and editing, rated blind, pairwise.

Rumeqo runs these models inside your team rooms. See what each one costs.

86 of 86 ranked models
Ranked models
rankmodelvendorcompositebenchmarkseloeloelo
1
GPT Image 2 (medium)openai/gpt-image-2:medium
OpenAI81.63 of 3 benchmarks138014631454
2
Muse Imagemeta/muse-image
Meta79.63 of 3 benchmarks128314051401
3
Seedream 5.0 Probytedance-seed/seedream-5-0-pro
ByteDance Seed78.23 of 3 benchmarks125713931414
4
Nano Banana 2 (Gemini 3.1 Flash Image)google/gemini-3.1-flash-image
Google75.13 of 3 benchmarks126313851365
5
Nano Banana Pro (Gemini 3 Pro Image Preview)google/gemini-3-pro-image-preview
Google74.13 of 3 benchmarks123213851368
6
Nano Banana Pro (Gemini 3 Pro Image) (2K)google/gemini-3-pro-image:2k
Google74.03 of 3 benchmarks124613891364
7
MAI-Image-2.5microsoft/mai-image-2.5
Microsoft73.52 of 3 benchmarks12561402
8
Reve 2.0reve/reve-2.0
Reve72.63 of 3 benchmarks127013581344
9
Reve 2.1reve/reve-2.1
Reve72.02 of 3 benchmarks13021374
10
Chatgpt Image High Fidelityopenai/chatgpt-image-high-fidelity
OpenAI71.02 of 3 benchmarks13901353
11
GPT Image 1.5 High Fidelityopenai/gpt-image-1.5-high-fidelity
OpenAI70.53 of 3 benchmarks123913701342
12
Grok Imagine Image Qualityx-ai/grok-imagine-image-quality
SpaceXAI70.32 of 3 benchmarks12281390
13
Uni-1.1 (max)luma/uni-1.1:max
Luma AI67.83 of 3 benchmarks118813341308
14
Qwen Image 3.0 Proqwen/qwen-image-3.0-pro
Qwen67.01 of 3 benchmarks1258
15
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)google/gemini-3.1-flash-lite-image
Google66.83 of 3 benchmarks125113141286
16
Uni-1.1luma/uni-1.1
Luma AI64.13 of 3 benchmarks118013121287
17
Qwen Image 2.0 Proqwen/qwen-image-2.0-pro
Qwen64.02 of 3 benchmarks11911303
18
Grok Imagine Imagex-ai/grok-imagine-image
xAI63.82 of 3 benchmarks11721330
19
Ideogram 4.0 Qualityideogram/ideogram-4.0-quality
Ideogram63.31 of 3 benchmarks1206
20
Mai Image 2microsoft/mai-image-2
Microsoft61.91 of 3 benchmarks1182
21
Cosmos3 Super Text2imagenvidia/cosmos3-super-text2image
NVIDIA61.51 of 3 benchmarks1181
22
Seedream 4.5bytedance-seed/seedream-4.5
ByteDance Seed60.13 of 3 benchmarks114613011291
23
Recraft V4.1 Utility Prorecraft/recraft-v4.1-utility-pro
Recraft60.11 of 3 benchmarks1169
24
Grok Imagine Image Prox-ai/grok-imagine-image-pro
xAI59.11 of 3 benchmarks1161
25
Hunyuan Image 3.0 Instructtencent/hunyuan-image-3.0-instruct
Tencent57.81 of 3 benchmarks1302
26
Reve v1.5reve/reve-v1.5
Reve57.81 of 3 benchmarks1154
27
Hunyuan Image 3.0tencent/hunyuan-image-3.0
Tencent57.31 of 3 benchmarks1151
28
FLUX.2 Maxblack-forest-labs/flux.2-max
Black Forest Labs56.53 of 3 benchmarks116212621250
29
Imagen Ultra 4.0 Generate 001google/imagen-ultra-4.0-generate-001
Google56.41 of 3 benchmarks1148
30
Seedream 5.0 Litebytedance/seedream-5.0-lite
ByteDance55.83 of 3 benchmarks113612931270
31
Wan2.7 Image Proalibaba/wan2.7-image-pro
Alibaba55.43 of 3 benchmarks110313021289
32
Nano Banana (Gemini 2.5 Flash Image)google/gemini-2.5-flash-image
Google55.03 of 3 benchmarks115012941234
33
Wan2.6 T2Ialibaba/wan2.6-t2i
Alibaba54.01 of 3 benchmarks1136
34
FLUX.2 Problack-forest-labs/flux.2-pro
Black Forest Labs53.93 of 3 benchmarks115512441238
35
Reve v1.1reve/reve-v1.1
Reve53.72 of 3 benchmarks12611261
36
Recraft V4.1 Prorecraft/recraft-v4.1-pro
Recraft53.61 of 3 benchmarks1130
37
Wan2.7 Imagealibaba/wan2.7-image
Alibaba53.13 of 3 benchmarks110013011280
38
Imagen 4.0 Generate 001google/imagen-4.0-generate-001
Google53.11 of 3 benchmarks1129
39
Qwen Image 2512qwen/qwen-image-2512
Qwen52.61 of 3 benchmarks1126
40
Kling Image O1kwaivgi/kling-image-o1
Kling AI52.52 of 3 benchmarks12511252
41
Krea 2 (medium)krea/krea-2:medium
Krea52.21 of 3 benchmarks1122
42
Hidream O1 Imagehidream/hidream-o1-image
HiDream51.71 of 3 benchmarks1118
43
Wan2.5 T2I Previewalibaba/wan2.5-t2i-preview
Alibaba51.31 of 3 benchmarks1117
44
FLUX.2 Flexblack-forest-labs/flux.2-flex
Black Forest Labs51.13 of 3 benchmarks115612251234
45
Qwen Image Editqwen/qwen-image-edit
Qwen50.31 of 3 benchmarks1241
46
Recraft V4recraft/recraft-v4
Recraft49.91 of 3 benchmarks1113
47
Seedream 4 (2K)bytedance/seedream-4:2k
ByteDance49.43 of 3 benchmarks114012711201
48
Krea 2 Turbokrea/krea-2-turbo
Krea48.91 of 3 benchmarks1110
49
Krea 2 Largekrea/krea-2-large
Krea48.01 of 3 benchmarks1105
50
Reve v1reve/reve-v1
Reve47.22 of 3 benchmarks12341222
51
Mai Image 1microsoft/mai-image-1
Microsoft46.61 of 3 benchmarks1093
52
Seedream 3bytedance/seedream-3
ByteDance46.21 of 3 benchmarks1082
53
Wan2.6 Imagealibaba/wan2.6-image
Alibaba46.02 of 3 benchmarks12301219
54
Z Image Turboalibaba/z-image-turbo
Alibaba45.71 of 3 benchmarks1081
55
Flux 2 Devblack-forest-labs/flux-2-dev
Black Forest Labs45.73 of 3 benchmarks114612241202
56
Qwen Image Prompt Extendqwen/qwen-image-prompt-extend
Qwen44.31 of 3 benchmarks1060
57
Reve v1.1 Fastreve/reve-v1.1-fast
Reve44.11 of 3 benchmarks1207
58
Qwen Image Edit 2511qwen/qwen-image-edit-2511
Qwen43.82 of 3 benchmarks12351173
59
Reve Edit Fastreve/reve-edit-fast
Reve43.51 of 3 benchmarks1198
60
Imagen 3.0 Generate 002google/imagen-3.0-generate-002
Google43.41 of 3 benchmarks1058
61
Qwen Imageqwen/qwen-image
Qwen42.91 of 3 benchmarks1057
62
Ideogram v3 Qualityideogram/ideogram-v3-quality
Ideogram42.51 of 3 benchmarks1049
63
Seedream 4 High Res Falbytedance/seedream-4-high-res-fal
ByteDance42.23 of 3 benchmarks111312171209
64
Photonluma/photon
Luma AI42.01 of 3 benchmarks1035
65
Runway Gen4runway/runway-gen4
Runway41.11 of 3 benchmarks1025
66
Flux 2 Klein 9Bblack-forest-labs/flux-2-klein-9b
Black Forest Labs40.83 of 3 benchmarks107012241213
67
Recraft V3recraft/recraft-v3
Recraft40.61 of 3 benchmarks1021
68
Flux 1.1 Problack-forest-labs/flux-1.1-pro
Black Forest Labs40.21 of 3 benchmarks1016
69
Lucid Originleonardo/lucid-origin
Leonardo AI39.71 of 3 benchmarks1013
70
Seedream 4 Falbytedance/seedream-4-fal
ByteDance39.53 of 3 benchmarks111612101156
71
Ideogram v2ideogram/ideogram-v2
Ideogram39.21 of 3 benchmarks1013
72
GLM Imagez-ai/glm-image
Z.ai38.81 of 3 benchmarks1010
73
Flux 1 Devblack-forest-labs/flux-1-dev
Black Forest Labs37.81 of 3 benchmarks969
74
DALL E 3openai/dall-e-3
OpenAI37.41 of 3 benchmarks968
75
Wan2.5 I2I Previewalibaba/wan2.5-i2i-preview
Alibaba36.82 of 3 benchmarks11811160
76
Stable Diffusion v35 Largestability-ai/stable-diffusion-v35-large
Stability AI36.41 of 3 benchmarks938
77
Step1x Editstepfun/step1x-edit
StepFun36.01 of 3 benchmarks998
78
GPT Image 1openai/gpt-image-1
OpenAI35.03 of 3 benchmarks111511391126
79
FLUX.2 Klein 4Bblack-forest-labs/flux.2-klein-4b
Black Forest Labs33.73 of 3 benchmarks103011881166
80
GPT Image 1 Miniopenai/gpt-image-1-mini
OpenAI33.03 of 3 benchmarks110911241111
81
Flux 1 Kontext (max)black-forest-labs/flux-1-kontext:max
Black Forest Labs32.03 of 3 benchmarks107411811059
82
Flux 1 Kontext Problack-forest-labs/flux-1-kontext-pro
Black Forest Labs31.33 of 3 benchmarks105911761061
83
Seededit 3.0bytedance/seededit-3.0
ByteDance30.22 of 3 benchmarks11391042
84
Bagelbytedance/bagel
ByteDance27.52 of 3 benchmarks8981026
85
Gemini 2.0 Flash Preview Image Generationgoogle/gemini-2.0-flash-preview-image-generation
Google24.93 of 3 benchmarks97510811058
86
Flux 1 Kontext Devblack-forest-labs/flux-1-kontext-dev
Black Forest Labs24.63 of 3 benchmarks94011491034
How this ranks

Every benchmark value becomes a percentile among the models that have it, so accuracy scores, Elo ratings and word error rates compare without hand-tuned scaling. Metrics where lower is better are inverted first. Raw values are never summed or averaged across benchmarks. A model's mean percentile is then shrunk toward the mean of the models that were broadly benchmarked, so a model tested twice cannot outrank a broadly tested one on two lucky results. Turning a data source off runs that same ranking code again in your browser over the sources you left on.

A model scored on fewer than 2 of the 3 ranked benchmarks in this category still ranks here, on the benchmarks it does have, and its row carries a partial coverage mark. On an equal score it sits under the model that earned the same number across more of the board.

Data sources

Turn a source off to drop every benchmark it feeds and rank the board again from what is left, in your browser. Turn them all off and the table has nothing to rank. Your choice follows you across the leaderboard pages.

  • LMArenaCC BY 4.0

    Arena ratings by LMArena, from the public leaderboard dataset.