# Qwen Image 2.0 Pro on the anchor ladder v3

Blind pairwise test on the photo-anchored quality scale. Ratings collected 2026-09-11 to 2026-09-22.

## Files

- `votes.csv` — every judgement, one row each: scene, category, the two sides compared, rater, rater role, and the choice on each question (the winning side's id, `tie`, or empty if unanswered).
- `results.json` — scores with 95% intervals, pairwise contrasts, results by category and the study design.
- `scenes.csv` — every scene: category, aspect ratio, prompt and, where a real photograph is the reference, its source (own photography or the Unsplash author and original).
- Images — separate archive, linked from the model page.

## Questions

- `prompt_adherence` — Prompt adherence: Which image follows the prompt better?
- `believability` — Believability: Which is more believable as a real photograph?
- `aesthetics` — Aesthetics: Which is more attractive?

## Participants

- `qwen-image-2-0-pro` — Qwen Image 2.0 Pro
- `gpt-image-2` — GPT Image 2
- `qwen-image-2` — Qwen Image 2.0
- `flux-2-klein` — FLUX.2 Klein
- `real-photograph` — Real photograph

## Read before reusing

- Prompt adherence is biased against photography. Prompts were written by describing the reference photographs, so a model executing text literally scores better than a photograph containing incidental detail. The absolute score is therefore built on believability alone.
- Tie share keeps rising: 25.4% in the original ladder study, 28.4% for Seedream 4.0, 34.1% here. Part of this is expected — Qwen Image 2.0 Pro sits between the two anchors, so more pairs are genuinely close. But the interface also changed between the first study and the later ones, so the two causes cannot be separated from these data alone.
- Rater agreement is low by the usual yardstick. Conclusions hold at the level of the whole corpus, not for a single scene.

## Licence

Data — CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Images — All rights reserved; provided for verification only.

## How to cite

> Everypixel research (2026). Qwen Image 2.0 Pro on the anchor ladder v3 [Data set]. Ratings collected 11–22 September 2026.
