SplotchScrapbook

Image-model bake-off

8 candidate image models were each handed the same 19 toddler drawings and the exact prompt the app sends. This page compares what each one cost, how long it took, and what it drew.

8 candidates19 drawings · 10 categories152 calls$8.73 spentAug 14, 2026
Archived comparison. This preserves the Aug 14, 2026 bake-off and its production recommendation at that time. View latest bake-off →
The pick

gpt-image-2 · low

Cheap, quick enough, and it never adds anything the child did not draw. No candidate refused a single drawing, the toy-sword probe included. One call failed with a provider error.

  • 2.0¢per image
  • 27 smedian wait
  • 0refusals

Why low, not medium

The two tiers never overlapped on time. Low took 23 to 35 seconds, medium 41 to 64. Medium's fastest picture was slower than low's slowest, every single time. For a two-year-old the wait is the whole experience, and low costs about a third as much.

What low gives up

Low paints a scribbled fill literally: a hard green-and-red seam on the half-and-half apple, and black bars across the cow whose back was scribbled over. Medium reads the same marks as intent and blends them. Two things make that acceptable. Low never adds a subject the child did not draw, and the literal reading may come from the prompt itself, which asks for "one flat, even area of that solid color". Try a prompt change before blaming the model.

Why the others lost

  • gpt-image-1-mini · low is the cheapest (1.0¢) and the fastest OpenAI tier (20 s), but it invents: three stray blue strokes came back as two dragons. That rules it out for an app whose promise is turning your drawing into a picture.
  • gpt-image-1.5 · medium looked like the fast-and-detailed middle on paper (25 s, ten times low's image tokens) but dropped the magic-brush reveal on the cat, shrank compositions, and banded the apple anyway, at 2.7 times the cost.
  • gpt-image-2 · high is not visibly better than medium and costs 3.3 times more for twice the wait.
  • The two Gemini models are by far the fastest, 7 to 8 seconds, but cost 1.9 and 3.3 times more per image than low, and Gemini 2.5 ignored the child's brown scribble on the cow entirely.

Slow is no longer broken

Generation moved to a background worker with a five-minute budget (ADR-0115), so no candidate can miss a deadline any more. The dashed 24-second line on the time bars is the limit the old one-request flow had. It stays on the chart because it is why that rework happened.

Where to look in the gallery

The cow separates the field on fidelity: the child scribbled brown on its back, Gemini 2.5 drew a black-and-white cow that ignored it, and low painted the scribble on as stripes. The outline-only cat is the cleanest test of whether a model keeps the child's own stroke colors or repaints the subject.

Read the numbers with their limits

One sample per cell, and six calls were running at once, so every time here is slower than a lone call would be. Compare candidates against each other, not against the clock.

Cost and speed

one row per candidate

Cost is what one generated image costs at list prices. Time is how long each call took, from request to response.

Sort
  1. gemini-2.5-flash-imagecurrent prod
    3.9¢per image
    $39 per 1,000 · 3.9× cheapest
    7.6 smedian
    p90 8.8 s · range 6.5–11 s
    19 / 19images
  2. gemini-3.1-flash-imagegemini candidate
    6.8¢per image
    $68 per 1,000 · 6.9× cheapest
    7.3 smedian
    p90 9.3 s · range 6.0–9.9 s
    19 / 19images
  3. gpt-image-2 · lowopenai candidate in production now
    2.0¢per image
    $20 per 1,000 · 2.1× cheapest
    27 smedian
    p90 33 s · range 23–35 s
    18 / 19images
    1 error
  4. gpt-image-2 · mediumopenai candidate
    5.8¢per image
    $58 per 1,000 · 5.9× cheapest
    51 smedian
    p90 61 s · range 41–64 s
    19 / 19images
  5. gpt-image-2 · highopenai candidate
    19¢per image
    $191 per 1,000 · 19.3× cheapest
    109 smedian
    p90 143 s · range 98–150 s
    19 / 19images
  6. gpt-image-1.5 · mediumopenai candidate
    5.6¢per image
    $56 per 1,000 · 5.6× cheapest
    25 smedian
    p90 31 s · range 18–37 s
    19 / 19images
  7. gpt-image-1-mini · lowopenai budget
    1.0¢per image
    $10 per 1,000 · cheapest
    20 smedian
    p90 27 s · range 16–33 s
    19 / 19images
  8. gpt-image-1-mini · mediumopenai budget
    1.8¢per image
    $19 per 1,000 · 1.9× cheapest
    25 smedian
    p90 32 s · range 21–41 s
    19 / 19images
medianp90 (9 in 10 calls finish by here)24 s, the limit of a single synchronous request

The dashed line is the app's original one-request limit: the server had to answer inside 24 s or Netlify cut it off (ADR-0063). Generation now runs in a background worker the app polls (ADR-0115), so a candidate past the line is slow, not unusable. Times were measured with 6 calls running at once, so they run slower than a single call on its own would.

All the numbers tokens, mean, min and max, file size
CandidateImagesImage tokens
(median)
$ per image$ per 1,000vs cheapest MeanMedianp90MinMaxRefusedErrorsFile size
gemini-2.5-flash-imagecurrent prod 19 / 19 1,290 $0.0389 $39 3.9× 7.8 s 7.6 s 8.8 s 6.5 s 11 s 0 0 1,160 KB
gemini-3.1-flash-imagegemini candidate 19 / 19 1,120 $0.0679 $68 6.9× 7.6 s 7.3 s 9.3 s 6.0 s 9.9 s 0 0 390 KB
gpt-image-2 · lowopenai candidate 18 / 19 158 $0.0204 $20 2.1× 28 s 27 s 33 s 23 s 35 s 0 1 1,445 KB
gpt-image-2 · mediumopenai candidate 19 / 19 1,372 $0.0581 $58 5.9× 51 s 51 s 61 s 41 s 64 s 0 0 1,633 KB
gpt-image-2 · highopenai candidate 19 / 19 5,488 $0.1906 $191 19.3× 116 s 109 s 143 s 98 s 150 s 0 0 1,554 KB
gpt-image-1.5 · mediumopenai candidate 19 / 19 1,584 $0.0559 $56 5.6× 26 s 25 s 31 s 18 s 37 s 0 0 1,531 KB
gpt-image-1-mini · lowopenai budget 19 / 19 408 $0.0099 $10 1.0× 21 s 20 s 27 s 16 s 33 s 0 0 1,587 KB
gpt-image-1-mini · mediumopenai budget 19 / 19 1,584 $0.0185 $19 1.9× 26 s 25 s 32 s 21 s 41 s 0 0 1,768 KB
Time by category mean seconds per drawing type
CategoryDrawingsgemini-2.5-flash-imagegemini-3.1-flash-imagegpt-image-2 · lowgpt-image-2 · mediumgpt-image-2 · highgpt-image-1.5 · mediumgpt-image-1-mini · lowgpt-image-1-mini · medium
Freehand scenes2873146126191823
Coloring page, magic brush2892656109312625
Coloring page, colored by hand2773047103292028
Coloring page, barely started2872457105282530
Filled drawings2982549117222533
Outlines only2982653124221923
Magic brush on blank paper2783156148281923
Night mode2873047110271928
Pretend-play probe17102446102232425
A few strokes, one color2882852106281924

Calls that returned no image

1 of 152

Gallery

every drawing, every candidate

Each drawing is shown first, followed by what every candidate made of it. Tap any result to swap it for the child's drawing and back, so you can see exactly what changed. Use the toolbar to hide candidates you have ruled out.

gpt-image-2
gpt-image-1.5
gpt-image-1-mini

How this was measured