Qwen Image 2.0 scores 64.5 on our photo-anchored scale, where a real photograph is 100; it is the middle rung of the ladder. In blind pairwise tests on 48 commercial scenes it clearly beat FLUX.2 Klein, fell clearly behind GPT Image 2 and the reference photographs, and showed no measurable difference from Seedream 4.0, leaning ahead on product shots and typography.
A real photograph is pinned at 100, and every model sits by how often it won blind comparisons — the further left, the less often. Scroll sideways to see the rest of the ruler.
Preferred over the real photograph in 2 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
Even in 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
87 · #3
87#3
Preferred over the real photograph in 2 of 3 blind comparisons
Even in 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
Preferred over the real photograph in 1 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
73 · #3
73#3
The real photograph was preferred in 2 of 3 blind comparisons
Preferred over the real photograph in 1 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
71 · #4
71#4
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
The real photograph was preferred in 1 of 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
69 · #3
69#3
The real photograph was preferred in 1 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
Preferred over the real photograph in 1 of 3 blind comparisons
The real photograph was preferred in 1 of 3 blind comparisons
46 · #4
46#4
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
The real photograph was preferred in 1 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
15 · #4
15#4
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 1 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
14 · #3
14#3
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
When to use
Same brief, side by side
Prompt used:A Black woman with long locs, bare shoulders and a strapless cream top stands among dense tropical foliage and looks calmly into the lens. She holds an amber glass dropper bottle with a gold cap up in front of her chest with both hands, fingertips lightly touching it. Broad green leaves and tall reed stems fill the background; a blurred leaf crosses the bottom of the frame.
Real photo
Qwen Image 2.0
The real photograph was preferred in 1 of 3 blind comparisons
The dependable middle of the ladder: a real step above the open-weights floor, level with Seedream 4.0, and a clear step below GPT Image 2. I would use it for everyday product and layout work and move the hero asset up a rung.
Strengths
Clearly ahead of FLUX.2 Klein on believability
Level with Seedream 4.0 overall
Leans ahead of Seedream 4.0 on product shots and typography
Weaknesses
Clearly behind GPT Image 2
Clearly behind real photographs on believability
Leans behind Seedream 4.0 on social and fashion scenes
Rater agreement is low by the usual yardstick (Fleiss' kappa 0.16–0.20). Half the pairs compare adjacent tiers, where “equal” is an honest answer. Conclusions hold at the level of the whole corpus, not for a single scene, pair or rater.
Prices are for one 1024×1024 image at standard settings, checked 25 September 2026; the price at the top of the page is the average of the providers listed. Workroom pricing is on the pricing page.
Qwen Image 2.0 is Alibaba's model, available through the Alibaba Cloud Model Studio API. Earlier Qwen-Image releases were published with open weights under Apache 2.0; check Alibaba's current terms for version 2.0.
Frequently asked questions
Not measurably. On believability the gap is +0.07 with a 95% interval from −0.14 to +0.28, which includes zero, and Qwen Image 2.0 was preferred in 51.8% of judgements with ties split. The difference is in the mix: Qwen leaned ahead on product shots, typography and street scenes, Seedream on social and fashion.
Qwen Image 2.0 is a rung of the ladder: its score was fitted when the scale was built and is frozen for later sessions. The published baseline gives the fitted scores of the rungs without score intervals, so the card shows the point value only rather than an invented range.
Clear. The believability strength gap is 0.83 in GPT Image 2's favour, and the dataset notes that the lower rungs separate reliably. The same gap was re-measured in the September session and the frozen value fell inside that session's interval. Across the three questions GPT Image 2 was preferred in 72.0% of judgements.
It was one of four rungs compared blind on 48 commercial scenes in August 2026, paired on every scene with FLUX.2 Klein, GPT Image 2 and the reference photographs, each pair judged three times on three questions. In September it was also paired with Seedream 4.0 on the same scenes.
No. Workroom is an AI router that gives access to third-party models, including Qwen Image 2.0, and is not affiliated with Alibaba. The judgement-level data behind this card is published by Everypixel research under CC BY 4.0, so every number on this page can be checked against the source dataset.
About Workroom
Workroom is an AI router: we give access to third-party models, including Qwen Image 2.0, and we are not affiliated with Alibaba. We have no stake in any single vendor winning; the router is only useful if it sends each brief to the right model. The judgement-level data behind this card is published by Everypixel research under CC BY 4.0 and linked below, so every number here can be checked independently.
How we tested
Between 21 and 25 August 2026 we compared four rungs of our quality ladder blind on 48 commercial scenes: FLUX.2 Klein, Qwen Image 2, GPT Image 2 and the real photographs each scene was written from. Every rung was paired with every other on every scene — 288 pairs and 863 judgements — and each pair was judged three times, without any indication of which image came from which model, on three questions: which image follows the prompt better, which is more believable as a real photograph, and which is more attractive. In September Qwen Image 2 was also paired blind with Seedream 4.0 on the same scenes: 432 more judgements.
How the scale is built
Pairwise comparisons are fitted with a Bradley–Terry model; intervals come from bootstrapping over scenes. The rungs of the scale are frozen from the baseline study, so every new model lands on the same scale. The score uses believability only:
score = 100 + (strength - strength_gold) / a
A real photograph is pinned at 100 by definition; slope 0.020835 strength per point; calibration points: floor, mid, gold. Category positions are fitted the same way on the 6 scenes of each category, with a real photograph pinned at 100 on every one of them.
Direct blind tests
The pairs Qwen Image 2.0 was judged against directly — where its place on the scale comes from.
Data: CC BY 4.0. Images: All rights reserved; ten Unsplash photographs under the Unsplash License.
Read these before reusing the data
Prompt adherence is biased against the photograph: each prompt was written by looking at the photograph, and a model executes text literally, while a photograph always contains incidental detail the description does not mention. Raters count that as a miss — which is why the absolute scale is built on believability alone.
The share of ties is inflated: during collection the “equal” option was pre-selected, so a criterion the rater never touched was recorded as a genuine tie. The defect was fixed on 28 August 2026, after this data was collected. Tier ordering survives it: dropping every tie only widens the gaps.
All images were shown at 2048 px on the long edge — native output size for the generative tiers; the photographs were downscaled from 30–45 megapixels. On screen the photographs are not softer than the generated images.
<blockquote cite="https://workroom.everypixel.com/models/qwen-image-2-0">
<p>Qwen Image 2.0 scores 64.5 on our photo-anchored scale, where a real photograph is 100; it is the middle rung of the ladder.</p>
<footer>— <a href="https://workroom.everypixel.com/models/qwen-image-2-0">Qwen Image 2.0 benchmark: blind pairwise test</a>, Everypixel Workroom, August 2026</footer>
</blockquote>