Skip to main content

Rank

#3

Score

65

Price

$0.035

Qwen Image 2.0 scores 64.5 on our photo-anchored scale, where a real photograph is 100; it is the middle rung of the ladder. In blind pairwise tests on 48 commercial scenes it clearly beat FLUX.2 Klein, fell clearly behind GPT Image 2 and the reference photographs, and showed no measurable difference from Seedream 4.0, leaning ahead on product shots and typography.

Vendor
Alibaba
Type
T2I
Modalities
→
Released
Feb 10, 2026
Max resolution
2K (2048×2048)

Where it stands

A real photograph is pinned at 100, and every model sits by how often it won blind comparisons — the further left, the less often. Scroll sideways to see the rest of the ruler.

  • 104 · #3

    • Even in 2 blind comparisons

    • Preferred over the real photograph in 2 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • Preferred over the real photograph in 2 of 3 blind comparisons

    • Even in 3 blind comparisons

    • Preferred over the real photograph in 2 of 3 blind comparisons

  • 87 · #3

    • Preferred over the real photograph in 2 of 3 blind comparisons

    • Even in 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • Preferred over the real photograph in 1 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

  • 73 · #3

    • The real photograph was preferred in 2 of 3 blind comparisons

    • Preferred over the real photograph in 1 of 3 blind comparisons

    • The real photograph was preferred in 3 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • Preferred over the real photograph in 2 of 3 blind comparisons

  • 71 · #4

    • The real photograph was preferred in 2 of 3 blind comparisons

    • The real photograph was preferred in 3 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • Preferred over the real photograph in 2 of 3 blind comparisons

    • The real photograph was preferred in 1 of 3 blind comparisons

    • Preferred over the real photograph in 2 of 3 blind comparisons

  • 69 · #3

    • The real photograph was preferred in 1 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • The real photograph was preferred in 3 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • Preferred over the real photograph in 1 of 3 blind comparisons

    • The real photograph was preferred in 1 of 3 blind comparisons

  • 46 · #4

    • The real photograph was preferred in 3 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • Preferred over the real photograph in 2 of 3 blind comparisons

    • The real photograph was preferred in 1 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

  • 15 · #4

    • The real photograph was preferred in 3 of 3 blind comparisons

    • The real photograph was preferred in 1 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • The real photograph was preferred in 3 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • Preferred over the real photograph in 2 of 3 blind comparisons

  • 14 · #3

    • The real photograph was preferred in 3 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • The real photograph was preferred in 3 of 3 blind comparisons

    • The real photograph was preferred in 2 of 3 blind comparisons

    • The real photograph was preferred in 3 of 3 blind comparisons

When to use

Same brief, side by side

Prompt used:A Black woman with long locs, bare shoulders and a strapless cream top stands among dense tropical foliage and looks calmly into the lens. She holds an amber glass dropper bottle with a gold cap up in front of her chest with both hands, fingertips lightly touching it. Broad green leaves and tall reed stems fill the background; a blurred leaf crosses the bottom of the frame.

The real photograph was preferred in 1 of 3 blind comparisons

View all 48 scenes →

Expert evaluation

Would use with caveats
The dependable middle of the ladder: a real step above the open-weights floor, level with Seedream 4.0, and a clear step below GPT Image 2. I would use it for everyday product and layout work and move the hero asset up a rung.

Strengths

  • Clearly ahead of FLUX.2 Klein on believability
  • Level with Seedream 4.0 overall
  • Leans ahead of Seedream 4.0 on product shots and typography

Weaknesses

  • Clearly behind GPT Image 2
  • Clearly behind real photographs on believability
  • Leans behind Seedream 4.0 on social and fashion scenes

Rater agreement is low by the usual yardstick (Fleiss' kappa 0.16–0.20). Half the pairs compare adjacent tiers, where “equal” is an honest answer. Conclusions hold at the level of the whole corpus, not for a single scene, pair or rater.

Dmitry Shironosov

Dmitry Shironosov

Head of AI Research, Everypixel Workroom

What it costs

API list prices for Qwen Image 2.0
ProviderPriceChecked
Alibaba Cloud Model StudioInternational deployment (Singapore)$0.035 / imgSep 25, 2026
fal.ai$0.035 / imgSep 25, 2026
Replicate$0.035 / imgSep 25, 2026
Runware$0.035 / imgSep 25, 2026

Prices are for one 1024×1024 image at standard settings, checked 25 September 2026; the price at the top of the page is the average of the providers listed. Workroom pricing is on the pricing page.

Workroom pricing →

License and terms

Qwen Image 2.0 is Alibaba's model, available through the Alibaba Cloud Model Studio API. Earlier Qwen-Image releases were published with open weights under Apache 2.0; check Alibaba's current terms for version 2.0.

Frequently asked questions

About Workroom

Workroom is an AI router: we give access to third-party models, including Qwen Image 2.0, and we are not affiliated with Alibaba. We have no stake in any single vendor winning; the router is only useful if it sends each brief to the right model. The judgement-level data behind this card is published by Everypixel research under CC BY 4.0 and linked below, so every number here can be checked independently.

How we tested

Between 21 and 25 August 2026 we compared four rungs of our quality ladder blind on 48 commercial scenes: FLUX.2 Klein, Qwen Image 2, GPT Image 2 and the real photographs each scene was written from. Every rung was paired with every other on every scene — 288 pairs and 863 judgements — and each pair was judged three times, without any indication of which image came from which model, on three questions: which image follows the prompt better, which is more believable as a real photograph, and which is more attractive. In September Qwen Image 2 was also paired blind with Seedream 4.0 on the same scenes: 432 more judgements.

How the scale is built

Pairwise comparisons are fitted with a Bradley–Terry model; intervals come from bootstrapping over scenes. The rungs of the scale are frozen from the baseline study, so every new model lands on the same scale. The score uses believability only:

score = 100 + (strength - strength_gold) / a

A real photograph is pinned at 100 by definition; slope 0.020835 strength per point; calibration points: floor, mid, gold. Category positions are fitted the same way on the 6 scenes of each category, with a real photograph pinned at 100 on every one of them.

Direct blind tests

The pairs Qwen Image 2.0 was judged against directly — where its place on the scale comes from.

Qwen Image 2.0 was preferred over FLUX.2 Klein in 69% of blind comparisons.

Gap (Bradley–Terry)
+0.85 (interval not published)
Preference
68.6% of blind comparisons preferred it (ties split, all questions, vote share)
Prompt adherence80 preferred · 47 equal · 17 other preferred
Believability88 preferred · 34 equal · 22 other preferred
Aesthetics73 preferred · 30 equal · 41 other preferred

144 blind comparisons, each pair shown 3 times

GPT Image 2 was preferred over Qwen Image 2.0 in 72% of blind comparisons.

Gap (Bradley–Terry)
−0.83 (interval not published)
Preference
28.0% of blind comparisons preferred it (ties split, all questions, vote share)
Prompt adherence23 preferred · 53 equal · 68 other preferred
Believability24 preferred · 27 equal · 93 other preferred
Aesthetics28 preferred · 12 equal · 104 other preferred

144 blind comparisons, each pair shown 3 times

vs Real photograph

The real photograph was preferred over Qwen Image 2.0 in 62% of blind comparisons.

Gap (Bradley–Terry)
−0.74 (interval not published)
Preference
37.6% of blind comparisons preferred it (ties split, all questions, vote share)
Prompt adherence36 preferred · 52 equal · 55 other preferred
Believability29 preferred · 38 equal · 76 other preferred
Aesthetics40 preferred · 23 equal · 80 other preferred

143 blind comparisons, each pair shown 3 times

Qwen Image 2.0 Pro was preferred over Qwen Image 2.0 in 47% of blind comparisons.

Gap (Bradley–Terry)
−0.45 [−0.72, −0.21]
Preference
24.8% of blind comparisons preferred it (wins over all judgements, believability only, vote share)
Prompt adherence78 preferred · 235 equal · 119 other preferred
Believability107 preferred · 122 equal · 203 other preferred
Aesthetics93 preferred · 67 equal · 272 other preferred

432 blind comparisons, each pair shown 3 times

A coin flip: 52% for Qwen Image 2.0, no measurable difference.

Gap (Bradley–Terry)
+0.07 [−0.14, +0.28]
Preference
51.8% of blind comparisons preferred it (ties split, all questions)
Prompt adherence120 preferred · 208 equal · 102 other preferred
Believability170 preferred · 133 equal · 129 other preferred
Aesthetics170 preferred · 78 equal · 183 other preferred

432 blind comparisons, each pair shown 3 times

Open data

Every image, every vote and every number behind the scores, so the conclusions can be checked independently.

Data: CC BY 4.0. Images: All rights reserved; ten Unsplash photographs under the Unsplash License.

Read these before reusing the data

  • Prompt adherence is biased against the photograph: each prompt was written by looking at the photograph, and a model executes text literally, while a photograph always contains incidental detail the description does not mention. Raters count that as a miss — which is why the absolute scale is built on believability alone.
  • The share of ties is inflated: during collection the “equal” option was pre-selected, so a criterion the rater never touched was recorded as a genuine tie. The defect was fixed on 28 August 2026, after this data was collected. Tier ordering survives it: dropping every tie only widens the gaps.
  • All images were shown at 2048 px on the long edge — native output size for the generative tiers; the photographs were downscaled from 30–45 megapixels. On screen the photographs are not softer than the generated images.

Read the full study →

Cite this page

<blockquote cite="https://workroom.everypixel.com/models/qwen-image-2-0">
  <p>Qwen Image 2.0 scores 64.5 on our photo-anchored scale, where a real photograph is 100; it is the middle rung of the ladder.</p>
  <footer>&mdash; <a href="https://workroom.everypixel.com/models/qwen-image-2-0">Qwen Image 2.0 benchmark: blind pairwise test</a>, Everypixel Workroom, August 2026</footer>
</blockquote>

Everypixel Workroom. (2026). Qwen Image 2.0 benchmark: blind pairwise test. https://workroom.everypixel.com/models/qwen-image-2-0