Skip to main content

Seedream 4.0

Rank

#4

Score

57

Price

$0.03

Seedream 4.0 scores 57 on our photo-anchored scale (95% interval 44–70), where a real photograph is 100. In 864 blind judgements on 48 commercial scenes it showed no measurable difference from Qwen Image 2.0: Seedream was preferred in 48.2% of judgements, ties split. GPT Image 2 was clearly ahead, preferred in 68.6%. By category, Seedream leaned ahead on social marketing and fashion, and behind on product shots, typography and street scenes.

Vendor
ByteDance
Family
Seedream
Type
T2I
Modalities
→
License
Proprietary
Released
Sep 9, 2025

Where it stands

A real photograph is pinned at 100, and every model sits by how often it won blind comparisons — the further left, the less often. Scroll sideways to see the rest of the ruler.

  • 83 · #4

  • 80 · #2

  • 79 · #3

  • 67 · #4

  • 57 · #3

  • 53 · #4

  • 16 · #3

  • −14 · #4

When to use

Same brief, side by side

Prompt used:A low-angle outdoor selfie of a young woman with dark curly hair in two afro puffs, large white round earrings and a white necklace, laughing widely with her arm stretched toward the camera. She wears a mint quilted bomber jacket over a navy and white striped top with a canvas tote strap on her shoulder, and holds a glass bottle of detox water with cucumber slices and lemon. Behind her a modern building facade with pink and mint vertical louvres, ornamental grasses and a pale concrete bench.

View all 48 scenes →

Expert evaluation

Would use with caveats
A second mid-tier generator rather than a step up: statistically level with Qwen Image 2.0 on believability, with a different category profile. Worth routing social and fashion briefs to; keep product and typography work on Qwen Image 2.0 or GPT Image 2.

Strengths

  • Level with Qwen Image 2.0 on believability across the corpus
  • Leans ahead on social-marketing and fashion scenes

Weaknesses

  • Clearly behind GPT Image 2
  • Behind Qwen Image 2.0 on product shots, typography and street scenes
  • No seed control, so results cannot be reproduced exactly
Dmitry Shironosov

Dmitry Shironosov

Head of AI Research, Everypixel Workroom

What it costs

API list prices for Seedream 4.0
ProviderPriceChecked
ByteDance (BytePlus ModelArk)$0.03 / imgSep 25, 2026
fal.ai$0.03 / imgSep 25, 2026
Together AI$0.03 per megapixel, 1 MP$0.03 / imgSep 25, 2026
Runware$0.03 / imgSep 25, 2026

Prices are for one 1024×1024 image at standard settings, checked 25 September 2026; the price at the top of the page is the average of the providers listed. Workroom pricing is on the pricing page.

Workroom pricing →

License and terms

Seedream 4.0 is a proprietary ByteDance model served through the BytePlus ModelArk API; its weights are not published. Terms for generated images are set by the service you generate them in.

Frequently asked questions

About Workroom

Workroom is an AI router: we give access to third-party models, including Seedream 4.0, and we are not affiliated with ByteDance. We have no stake in any single vendor winning; the router is only useful if it sends each brief to the right model. The judgement-level data behind this card is published by Everypixel research under CC BY 4.0 and linked below, so every number here can be checked independently.

How we tested

Between 6 and 11 September 2026 we compared Seedream 4.0 blind against two fixed rungs of our quality ladder, Qwen Image 2 and GPT Image 2, on the same 48 commercial scenes the ladder was built on — three independent generations per scene, 288 pairs, 864 judgements. Every pair was judged three times, without any indication of which image came from which model, on three questions: which image follows the prompt better, which is more believable as a real photograph, and which is more attractive.

How the scale is built

Pairwise comparisons are fitted with a Bradley–Terry model; intervals come from bootstrapping over scenes. The rungs of the scale are frozen from the baseline study, so every new model lands on the same scale. The score uses believability only:

score = 100 + (strength − strength_reference) / a

model strength fitted against frozen anchor strengths from the v3 baseline; the calibration line is pinned at the photograph = 100 Category positions are fitted the same way on the 6 scenes of each category, with a real photograph pinned at 100 on every one of them.

The scale reproduced

The frozen gap between the two rungs was 0.83; this session measured 0.71 with a 95% interval of 0.45 to 1.02. The frozen value falls inside that interval.

Direct blind tests

The pairs Seedream 4.0 was judged against directly — where its place on the scale comes from.

A coin flip: 48% for Seedream 4.0, no measurable difference.

Gap (Bradley–Terry)
−0.07 [−0.28, +0.14]
Preference
48.2% of blind comparisons preferred it (ties split, all questions)
Prompt adherence102 preferred · 208 equal · 120 other preferred
Believability129 preferred · 133 equal · 170 other preferred
Aesthetics183 preferred · 78 equal · 170 other preferred

432 blind comparisons, each pair shown 3 times

GPT Image 2 was preferred over Seedream 4.0 in 69% of blind comparisons.

Gap (Bradley–Terry)
−0.78 [−1.03, −0.55]
Preference
31.4% of blind comparisons preferred it (ties split, all questions)
Prompt adherence74 preferred · 171 equal · 187 other preferred
Believability73 preferred · 94 equal · 265 other preferred
Aesthetics101 preferred · 52 equal · 279 other preferred

432 blind comparisons, each pair shown 3 times

Open data

Every image, every vote and every number behind the scores, so the conclusions can be checked independently.

Data: CC BY 4.0. Images: All rights reserved; provided for verification only.

Read these before reusing the data

  • The model ignores the requested output size: Seedream returned 2352 px on the long edge where 2048 was asked. Output was downscaled to match the ladder; it was never upscaled.
  • No seed control: Seedream does not accept a seed, so the three generations per scene are independent but not reproducible. This measures average ability rather than a best pick.
  • Prompt adherence is biased against photography, which is why the absolute score is built on believability alone.
  • Tie share is 28.4%, higher than the 25.4% of the original ladder study. This session pairs two models that are genuinely indistinguishable, where “equal” is an honest answer.

Cite this page

<blockquote cite="https://workroom.everypixel.com/models/seedream-4-0">
  <p>Seedream 4.0 scores 57 on our photo-anchored scale (95% interval 44–70), where a real photograph is 100.</p>
  <footer>&mdash; <a href="https://workroom.everypixel.com/models/seedream-4-0">Seedream 4.0 benchmark: blind pairwise test</a>, Everypixel Workroom, September 2026</footer>
</blockquote>

Everypixel Workroom. (2026). Seedream 4.0 benchmark: blind pairwise test. https://workroom.everypixel.com/models/seedream-4-0