# Anchor comparison of generative image models, v3

Blind pairwise test on the photo-anchored quality scale. Ratings collected 2026-08-21 to 2026-08-25.

## Files

- `votes.csv` — every judgement, one row each: scene, category, the two sides compared, rater, rater role, and the choice on each question (the winning side's id, `tie`, or empty if unanswered).
- `results.json` — scores with 95% intervals, pairwise contrasts, results by category and the study design.
- `scenes.csv` — every scene: category, aspect ratio, prompt and, where a real photograph is the reference, its source (own photography or the Unsplash author and original).
- Images — separate archive, linked from the model page.

## Questions

- `prompt_adherence` — Prompt adherence: Which image follows the prompt better?
- `believability` — Believability: Which is more believable as a real photograph?
- `aesthetics` — Aesthetics: Which is more attractive?

## Participants

- `flux-2-klein` — FLUX.2 Klein
- `qwen-image-2` — Qwen Image 2.0
- `gpt-image-2` — GPT Image 2
- `real-photograph` — Real photograph

## Read before reusing

- Prompt adherence is biased against the photograph: each prompt was written by looking at the photograph, and a model executes text literally, while a photograph always contains incidental detail the description does not mention. Raters count that as a miss — which is why the absolute scale is built on believability alone.
- The share of ties is inflated: during collection the “equal” option was pre-selected, so a criterion the rater never touched was recorded as a genuine tie. The defect was fixed on 28 August 2026, after this data was collected. Tier ordering survives it: dropping every tie only widens the gaps.
- All images were shown at 2048 px on the long edge — native output size for the generative tiers; the photographs were downscaled from 30–45 megapixels. On screen the photographs are not softer than the generated images.

## Licence

Data — CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Images — All rights reserved; ten Unsplash photographs under the Unsplash License.

## How to cite

> Everypixel research (2026). Anchor comparison of generative image models, v3 [Data set]. Ratings collected 21–25 August 2026.
