Skip to main content

Qwen Image 2.0 Pro

Rank

#2

Score

80

Price

$0.076

Qwen Image 2.0 Pro scores 80 on our photo-anchored scale (95% interval 69–94), where a real photograph is 100. In blind pairwise tests on 48 commercial scenes it was clearly ahead of Qwen Image 2.0, the non-Pro version, and clearly behind GPT Image 2. Its largest lead over Qwen Image 2.0 came on people scenes and on-image typography.

Vendor
Alibaba
Type
T2I
Modalities
→
Released
Mar 3, 2026
Max resolution
2K (2048×2048)

Where it stands

A real photograph is pinned at 100, and every model sits by how often it won blind comparisons — the further left, the less often. Scroll sideways to see the rest of the ruler.

  • 115 · #1

  • 96 · #2

  • 88 · #2

  • 82 · #2

  • 67 · #4

  • 62 · #2

  • 57 · #2

  • 51 · #2

When to use

Same brief, side by side

Prompt used:Two young women working together at a home-office desk. The one on the left, in a light blue pinstripe shirt with her dark hair tied back, leans in and points with her index finger at a large silver all-in-one monitor; the screen shows a photo-editing application with dark grey panels, a colour picker and a warm candlelit still-life open in the canvas. The second woman, seen from behind over her shoulder in the right foreground, out of focus, holds a stylus over a graphics tablet. On the wooden desktop: a wide black keyboard, a mouse, a camera body, a lens cap, potted plants on the windowsill, an angled desk lamp glowing warm, a window with white venetian blinds behind them.

View all 48 scenes →

Expert evaluation

Would use with caveats
A real generational step: the Pro version sits between its own predecessor and GPT Image 2, exactly where the ladder bracketed it. I would move everyday Qwen work to it, especially people and layouts with text, and still finish the most demanding hero asset a rung higher.

Strengths

  • Clearly ahead of Qwen Image 2.0 on believability
  • Widest lead over Qwen Image 2.0 on people scenes and typography
  • Honoured the requested size and seed on every generation

Weaknesses

  • Clearly behind GPT Image 2
  • Wide score interval, 69 to 94
  • Lowest category scores on typography and people scenes
Dmitry Shironosov

Dmitry Shironosov

Head of AI Research, Everypixel Workroom

What it costs

API list prices for Qwen Image 2.0 Pro
ProviderPriceChecked
Alibaba Cloud Model StudioInternational deployment (Singapore)$0.075 / imgSep 25, 2026
fal.ai$0.075 / imgSep 25, 2026
Runware$0.075 / imgSep 25, 2026
Together AI$0.08 / imgSep 25, 2026

Prices are for one 1024×1024 image at standard settings, checked 25 September 2026; the price at the top of the page is the average of the providers listed. Workroom pricing is on the pricing page.

Workroom pricing →

License and terms

Qwen Image 2.0 Pro is Alibaba's model, available through the Alibaba Cloud Model Studio API. Terms for generated images are set by the service you generate them in.

Frequently asked questions

About Workroom

Workroom is an AI router: we give access to third-party models, including Qwen Image 2.0 Pro, and we are not affiliated with Alibaba. We have no stake in any single vendor winning; the router is only useful if it sends each brief to the right model. The judgement-level data behind this card is published by Everypixel research under CC BY 4.0 and linked below, so every number here can be checked independently.

How we tested

Between 11 and 22 September 2026 Qwen Image 2.0 Pro was placed on our quality ladder. It got the same 48 commercial scenes and prompts as the ladder, three generations per scene, and each generation was paired blind with two fixed rungs, Qwen Image 2.0 and GPT Image 2: 288 pairs and 863 judgements. Each pair was judged three times on three questions: which image follows the prompt better, which is more believable as a real photograph, and which is more attractive.

How the scale is built

Pairwise comparisons are fitted with a Bradley–Terry model; intervals come from bootstrapping over scenes. The rungs of the scale are frozen from the baseline study, so every new model lands on the same scale. The score uses believability only:

score = 100 + (strength − strength_reference) / a

model strength fitted against frozen anchor strengths from the v3 baseline; the calibration line is pinned at the photograph = 100 Category positions are fitted the same way on the 6 scenes of each category, with a real photograph pinned at 100 on every one of them.

The scale reproduced

The frozen gap between the two rungs was 0.83; this session measured 1.09 with a 95% interval of 0.78 to 1.39. The frozen value falls inside that interval.

Direct blind tests

The pairs Qwen Image 2.0 Pro was judged against directly — where its place on the scale comes from.

Qwen Image 2.0 Pro was preferred over Qwen Image 2.0 in 47% of blind comparisons.

Gap (Bradley–Terry)
+0.45 [+0.21, +0.72]
Preference
47.0% of blind comparisons preferred it (wins over all judgements, believability only)
Prompt adherence119 preferred · 235 equal · 78 other preferred
Believability203 preferred · 122 equal · 107 other preferred
Aesthetics272 preferred · 67 equal · 93 other preferred

432 blind comparisons, each pair shown 3 times

GPT Image 2 was preferred over Qwen Image 2.0 Pro in 51% of blind comparisons.

Gap (Bradley–Terry)
−0.64 [−0.86, −0.42]
Preference
19.7% of blind comparisons preferred it (wins over all judgements, believability only)
Prompt adherence59 preferred · 246 equal · 126 other preferred
Believability85 preferred · 128 equal · 218 other preferred
Aesthetics112 preferred · 84 equal · 235 other preferred

431 blind comparisons, each pair shown 3 times

Open data

Every image, every vote and every number behind the scores, so the conclusions can be checked independently.

Data: CC BY 4.0. Images: All rights reserved; provided for verification only.

Read these before reusing the data

  • Prompt adherence is biased against photography. Prompts were written by describing the reference photographs, so a model executing text literally scores better than a photograph containing incidental detail. The absolute score is therefore built on believability alone.
  • Tie share keeps rising: 25.4% in the original ladder study, 28.4% for Seedream 4.0, 34.1% here. Part of this is expected — Qwen Image 2.0 Pro sits between the two anchors, so more pairs are genuinely close. But the interface also changed between the first study and the later ones, so the two causes cannot be separated from these data alone.
  • Rater agreement is low by the usual yardstick. Conclusions hold at the level of the whole corpus, not for a single scene.

Cite this page

<blockquote cite="https://workroom.everypixel.com/models/qwen-image-2-0-pro">
  <p>Qwen Image 2.0 Pro scores 80 on our photo-anchored scale (95% interval 69–94), where a real photograph is 100.</p>
  <footer>&mdash; <a href="https://workroom.everypixel.com/models/qwen-image-2-0-pro">Qwen Image 2.0 Pro benchmark: blind pairwise test</a>, Everypixel Workroom, September 2026</footer>
</blockquote>

Everypixel Workroom. (2026). Qwen Image 2.0 Pro benchmark: blind pairwise test. https://workroom.everypixel.com/models/qwen-image-2-0-pro