FLUX.2 Klein scores 23.9 on our photo-anchored scale, where a real photograph is 100; it is the lowest rung of the ladder. In blind pairwise tests on 48 commercial scenes it fell clearly behind Qwen Image 2.0, GPT Image 2 and the reference photographs, and it was the model furthest from the photographs. Use it for drafts and exploration rather than final photoreal images.
A real photograph is pinned at 100, and every model sits by how often it won blind comparisons — the further left, the less often. Scroll sideways to see the rest of the ruler.
The real photograph was preferred in 2 of 3 blind comparisons
Even in 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 1 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
Preferred over the real photograph in 3 of 3 blind comparisons
49 · #5
49#5
Preferred over the real photograph in 2 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
47 · #5
47#5
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
43 · #5
43#5
Even in 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
Preferred over the real photograph in 2 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
34 · #5
34#5
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 1 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
Preferred over the real photograph in 1 of 3 blind comparisons
−21 · #5
−21#5
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 1 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
The real photograph was preferred in 2 of 3 blind comparisons
−41 · #5
−41#5
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 1 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
−56 · #5
−56#5
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
The real photograph was preferred in 3 of 3 blind comparisons
When to use
Same brief, side by side
Prompt used:A Black woman with long locs, bare shoulders and a strapless cream top stands among dense tropical foliage and looks calmly into the lens. She holds an amber glass dropper bottle with a gold cap up in front of her chest with both hands, fingertips lightly touching it. Broad green leaves and tall reed stems fill the background; a blurred leaf crosses the bottom of the frame.
The floor of the ladder, and useful as one: it shows how far the stronger rungs move a brief. I would sketch directions in it, but nothing that has to read as a photograph should leave it unfinished.
Strengths
Open-weights model at the floor of the ladder
A clear baseline that shows how far stronger models move a brief
Weaknesses
Clearly behind Qwen Image 2.0
Clearly behind GPT Image 2
Furthest of all tested models from real photographs
Rater agreement is low by the usual yardstick (Fleiss' kappa 0.16–0.20). Half the pairs compare adjacent tiers, where “equal” is an honest answer. Conclusions hold at the level of the whole corpus, not for a single scene, pair or rater.
Prices are for one 1024×1024 image at standard settings, checked 25 September 2026; the price at the top of the page is the average of the providers listed. Workroom pricing is on the pricing page.
FLUX.2 Klein comes with open weights from Black Forest Labs in two sizes. The larger one, tested here, is released under the FLUX Non-Commercial License; the smaller one is released under Apache 2.0.
Frequently asked questions
Clearly behind. The believability strength gap is 0.85 in Qwen Image 2.0's favour, and the dataset notes that the three lower rungs separate reliably. Across the three questions Qwen Image 2.0 was preferred in 68.6% of judgements, counted from the published votes with ties split.
The ladder needs a clearly weaker rung to fix the slope of the scale. FLUX.2 Klein was chosen as that floor and in blind tests it separated reliably from Qwen Image 2 and the photographs. Together with them it is one of the three calibration points the scale is fitted on.
No. The scale measures one thing: how believable an image is as a real photograph, with a photograph fixed at 100. A low score means that, judged blind, the other image was picked as more photographic more often. Drafts, concepts and non-photographic work are not what this scale measures.
It was one of four rungs compared blind on 48 commercial scenes in August 2026, paired on every scene with Qwen Image 2, GPT Image 2 and the reference photographs: 288 pairs and 863 judgements in the study, each pair judged three times on prompt adherence, believability and aesthetics.
No. Workroom is an AI router that gives access to third-party models, including FLUX.2 Klein, and is not affiliated with Black Forest Labs. The judgement-level data behind this card is published by Everypixel research under CC BY 4.0, so every number on this page can be checked against the source dataset.
About Workroom
Workroom is an AI router: we give access to third-party models, including FLUX.2 Klein, and we are not affiliated with Black Forest Labs. We have no stake in any single vendor winning; the router is only useful if it sends each brief to the right model. The judgement-level data behind this card is published by Everypixel research under CC BY 4.0 and linked below, so every number here can be checked independently.
How we tested
Between 21 and 25 August 2026 we compared four rungs of our quality ladder blind on 48 commercial scenes: FLUX.2 Klein, Qwen Image 2, GPT Image 2 and the real photographs each scene was written from. Every rung was paired with every other on every scene — 288 pairs and 863 judgements — and each pair was judged three times, without any indication of which image came from which model, on three questions: which image follows the prompt better, which is more believable as a real photograph, and which is more attractive.
How the scale is built
Pairwise comparisons are fitted with a Bradley–Terry model; intervals come from bootstrapping over scenes. The rungs of the scale are frozen from the baseline study, so every new model lands on the same scale. The score uses believability only:
score = 100 + (strength - strength_gold) / a
A real photograph is pinned at 100 by definition; slope 0.020835 strength per point; calibration points: floor, mid, gold. Category positions are fitted the same way on the 6 scenes of each category, with a real photograph pinned at 100 on every one of them.
Direct blind tests
The pairs FLUX.2 Klein was judged against directly — where its place on the scale comes from.
Data: CC BY 4.0. Images: All rights reserved; ten Unsplash photographs under the Unsplash License.
Read these before reusing the data
Prompt adherence is biased against the photograph: each prompt was written by looking at the photograph, and a model executes text literally, while a photograph always contains incidental detail the description does not mention. Raters count that as a miss — which is why the absolute scale is built on believability alone.
The share of ties is inflated: during collection the “equal” option was pre-selected, so a criterion the rater never touched was recorded as a genuine tie. The defect was fixed on 28 August 2026, after this data was collected. Tier ordering survives it: dropping every tie only widens the gaps.
All images were shown at 2048 px on the long edge — native output size for the generative tiers; the photographs were downscaled from 30–45 megapixels. On screen the photographs are not softer than the generated images.
<blockquote cite="https://workroom.everypixel.com/models/flux-2-klein">
<p>FLUX.2 Klein scores 23.9 on our photo-anchored scale, where a real photograph is 100; it is the lowest rung of the ladder.</p>
<footer>— <a href="https://workroom.everypixel.com/models/flux-2-klein">FLUX.2 Klein benchmark: blind pairwise test</a>, Everypixel Workroom, August 2026</footer>
</blockquote>
Everypixel Workroom. (2026). FLUX.2 Klein benchmark: blind pairwise test. https://workroom.everypixel.com/models/flux-2-klein