GPT Image 2 scores 104.4 on our photo-anchored scale, where a real photograph is fixed at 100. That sits above the scale's ceiling but is not a claim of parity: in blind tests on 48 commercial scenes its believability gap to the photographs, +0.09 with a 95% interval from −0.16 to +0.33, was not separated. It clearly beat Qwen Image 2.0, FLUX.2 Klein and Seedream 4.0.
A real photograph is pinned at 100, and every model sits by how often it won blind comparisons — the further left, the less often. Scroll sideways to see the rest of the ruler.
Preferred over the real photograph in 1 of 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
Even in 3 blind comparisons
Even in 3 blind comparisons
Even in 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
114 · #2
114#2
The real photograph was preferred in 2 of 3 blind comparisons
Even in 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
Preferred over the real photograph in 3 of 3 blind comparisons
The real photograph was preferred in 1 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
114 · #1
114#1
Even in 3 blind comparisons
Even in 3 blind comparisons
Even in 3 blind comparisons
Preferred over the real photograph in 3 of 3 blind comparisons
The real photograph was preferred in 1 of 3 blind comparisons
Even in 3 blind comparisons
112 · #1
112#1
Preferred over the real photograph in 2 of 3 blind comparisons
Even in 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
Even in 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
Preferred over the real photograph in 3 of 3 blind comparisons
112 · #1
112#1
Preferred over the real photograph in 3 of 3 blind comparisons
Preferred over the real photograph in 3 of 3 blind comparisons
The real photograph was preferred in 1 of 3 blind comparisons
Preferred over the real photograph in 1 of 3 blind comparisons
Even in 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
91 · #1
91#1
Preferred over the real photograph in 1 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 1 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
85 · #1
85#1
The real photograph was preferred in 3 of 3 blind comparisons
Even in 3 blind comparisons
Even in 3 blind comparisons
Even in 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
Preferred over the real photograph in 1 of 3 blind comparisons
80 · #1
80#1
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
Even in 3 blind comparisons
Even in 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
Even in 3 blind comparisons
When to use
Same brief, side by side
Prompt used:A Black woman with long locs, bare shoulders and a strapless cream top stands among dense tropical foliage and looks calmly into the lens. She holds an amber glass dropper bottle with a gold cap up in front of her chest with both hands, fingertips lightly touching it. Broad green leaves and tall reed stems fill the background; a blurred leaf crosses the bottom of the frame.
Real photo
GPT Image 2
Preferred over the real photograph in 1 of 3 blind comparisons
The top rung we have: in blind tests its images were not reliably told apart from real photographs on believability. That is the ceiling of this scale, not proof of equality, but it is the model I would finish a hero image in.
Strengths
Not separated from real photographs on believability
Clearly ahead of every other model it was paired with
Widest lead over the open-weights floor, FLUX.2 Klein
Weaknesses
No measured lead over photographs: the score above 100 is the scale's ceiling
Close to even with photographs on believability
Rater agreement is low by the usual yardstick (Fleiss' kappa 0.16–0.20). Half the pairs compare adjacent tiers, where “equal” is an honest answer. Conclusions hold at the level of the whole corpus, not for a single scene, pair or rater.
Prices are for one medium-quality 1024×1024 image, checked 25 September 2026; the price at the top of the page is the average of the providers listed. OpenAI itself prices GPT Image 2 by tokens rather than per image, so it is not in the table. Workroom pricing is on the pricing page.
GPT Image 2 is a proprietary OpenAI model available through the OpenAI API; its weights are not published. Terms for generated images are set by the service you generate them in.
Frequently asked questions
This data makes no such claim. Its believability gap to the reference photographs is +0.09 with a 95% interval from −0.16 to +0.33, which includes zero, so the two are not separated. The 104.4 score reflects the scale's ceiling: a photograph is fixed at 100 and the scale has no room above it to measure a lead.
A calibration point has to be separated from its neighbour, otherwise the slope of the scale becomes arbitrary. Because GPT Image 2 was not separated from the photographs, the scale was fitted on FLUX.2 Klein, Qwen Image 2 and the photographs, and GPT Image 2's score was then read off that line.
Clearly ahead. In the September session the believability gap was +0.78 with a 95% interval from +0.55 to +1.03, which does not include zero, and GPT Image 2 was preferred in 68.6% of 864 judgements with ties split. On the photo-anchored scale that is 104 against 57.
It was one of four rungs compared blind on 48 commercial scenes in August 2026, paired on every scene with FLUX.2 Klein, Qwen Image 2 and the reference photographs, each pair judged three times on three questions. In September it was also paired with Seedream 4.0 on the same scenes.
No. Workroom is an AI router that gives access to third-party models, including GPT Image 2, and is not affiliated with OpenAI. The judgement-level data behind this card is published by Everypixel research under CC BY 4.0, so every number on this page can be checked against the source dataset.
About Workroom
Workroom is an AI router: we give access to third-party models, including GPT Image 2, and we are not affiliated with OpenAI. We have no stake in any single vendor winning; the router is only useful if it sends each brief to the right model. The judgement-level data behind this card is published by Everypixel research under CC BY 4.0 and linked below, so every number here can be checked independently.
How we tested
Between 21 and 25 August 2026 we compared four rungs of our quality ladder blind on 48 commercial scenes: FLUX.2 Klein, Qwen Image 2, GPT Image 2 and the real photographs each scene was written from. Every rung was paired with every other on every scene — 288 pairs and 863 judgements — and each pair was judged three times, without any indication of which image came from which model, on three questions: which image follows the prompt better, which is more believable as a real photograph, and which is more attractive. In September GPT Image 2 was also paired blind with Seedream 4.0 on the same scenes: 432 more judgements.
How the scale is built
Pairwise comparisons are fitted with a Bradley–Terry model; intervals come from bootstrapping over scenes. The rungs of the scale are frozen from the baseline study, so every new model lands on the same scale. The score uses believability only:
score = 100 + (strength - strength_gold) / a
A real photograph is pinned at 100 by definition; slope 0.020835 strength per point; calibration points: floor, mid, gold. Category positions are fitted the same way on the 6 scenes of each category, with a real photograph pinned at 100 on every one of them.
Direct blind tests
The pairs GPT Image 2 was judged against directly — where its place on the scale comes from.
Data: CC BY 4.0. Images: All rights reserved; ten Unsplash photographs under the Unsplash License.
Read these before reusing the data
Prompt adherence is biased against the photograph: each prompt was written by looking at the photograph, and a model executes text literally, while a photograph always contains incidental detail the description does not mention. Raters count that as a miss — which is why the absolute scale is built on believability alone.
The share of ties is inflated: during collection the “equal” option was pre-selected, so a criterion the rater never touched was recorded as a genuine tie. The defect was fixed on 28 August 2026, after this data was collected. Tier ordering survives it: dropping every tie only widens the gaps.
All images were shown at 2048 px on the long edge — native output size for the generative tiers; the photographs were downscaled from 30–45 megapixels. On screen the photographs are not softer than the generated images.
<blockquote cite="https://workroom.everypixel.com/models/gpt-image-2">
<p>GPT Image 2 scores 104.4 on our photo-anchored scale, where a real photograph is fixed at 100.</p>
<footer>— <a href="https://workroom.everypixel.com/models/gpt-image-2">GPT Image 2 benchmark: blind pairwise test</a>, Everypixel Workroom, August 2026</footer>
</blockquote>