SplotchScrapbook

Image-model bake-off

5 candidate image models were each handed the same 14 toddler drawings and the exact prompt the app sends. This page compares what each one cost, how long it took, and what it drew.

5 candidates14 drawings · 14 categories140 calls · 2 per cell$4.85 spentSep 10, 2026

Flare low is the leading challenger

A modest speed improvement is worth validating. The screen does not establish a universal visual winner or intrinsic cost savings.

Same winning prompt, fresh baseline

Fourteen synthetic drawings, one per category, two samples per configuration: 140 finished comparison images. Every call uses the shipped “Paint directly over” prompt (ADR-0118), with GPT-5.6 Sol at medium reasoning and the existing safety instruction.

What the averages say

Flare low: 23.0 s median versus 24.5 s for the baseline; p90 24.0 s versus 26.7 s. The broad run allowed three calls in flight. Mean layout scores: GPT Image 2 low 83.0; Flare low 85.5; Flare medium 86.8; Sunburst low 84.5; Sunburst medium 85.1.

The cost difference is mostly cache

Recorded mean cost is 3.16¢ for Flare low versus 3.53¢ for GPT Image 2 low. Removing the orchestrator cache discount puts every low tier near 3.94¢; the image tool itself averages about 1.52¢ on all three low-tier models. Do not interpret this run as a cheaper image-token price.

Look at the tradeoffs

The apple shows better blending in some medium-tier images. The colored balloon shows newer models sometimes preserving the red mark as a patch instead of coloring the balloon. The black scribbles show color changes despite high layout scores. Higher alignment scores do not automatically mean a better illustration.

Reliability and scope

The first pass produced 139 images and one baseline API 500. One targeted retry filled the missing image; the original failure is preserved in the run evidence. There were no refusals. The dark-paper row uses the base prompt without the Night Mode suffix. The blocked-content safety corpus and style suffixes were not tested.

Sequential confirmation

The separate counterbalanced check made 16 sequential calls over four inputs (eight per model). GPT Image 2 low: 24.2 s median, 31.8 s p90. Flare low: 21.8 s median, 23.0 s p90. Model order alternated AB/BA within each input and which model started alternated between inputs. This is a small confirmation sample, not a statistically powered latency study.

Before a production switch, validate the finalist with the actual style/night requests and the full safety suite. The measured visual differences do not justify skipping those checks.

Cost and speed

one row per candidate

Cost is what one generated image costs at list prices. Time is how long each call took, from request to response.

Sort
  1. 2 · lowprevious prod
    3.5¢per image
    $35 per 1,000 · 1.1× cheapest
    24 smedian
    p90 27 s · range 18–35 s
    28 / 28images
  2. flare · lowcurrent prod in production now
    3.2¢per image
    $32 per 1,000 · 1.0× cheapest
    23 smedian
    p90 24 s · range 17–27 s
    28 / 28images
  3. flare · medopenai candidate
    3.7¢per image
    $37 per 1,000 · 1.2× cheapest
    24 smedian
    p90 26 s · range 18–30 s
    28 / 28images
  4. sunburst · lowopenai candidate
    3.2¢per image
    $32 per 1,000 · cheapest
    26 smedian
    p90 29 s · range 19–31 s
    28 / 28images
  5. sunburst · medopenai candidate
    3.7¢per image
    $37 per 1,000 · 1.2× cheapest
    28 smedian
    p90 31 s · range 21–32 s
    28 / 28images
medianp90 (9 in 10 calls finish by here)24 s, the limit of a single synchronous request

The dashed line is the app's original one-request limit: the server had to answer inside 24 s or Netlify cut it off (ADR-0063). Generation now runs in a background worker the app polls (ADR-0115), so a candidate past the line is slow, not unusable. Times were measured with 3 calls running at once, so they run slower than a single call on its own would.

All the numbers tokens, mean, min and max, file size
CandidateImagesImage tokens
(median)
$ per image$ per 1,000vs cheapest MeanMedianp90MinMaxRefusedErrorsFile size
2 · lowprevious prod 28 / 28 158 $0.0353 $35 1.1× 24 s 24 s 27 s 18 s 35 s 0 0 1,505 KB
flare · lowcurrent prod 28 / 28 158 $0.0316 $32 1.0× 22 s 23 s 24 s 17 s 27 s 0 0 1,439 KB
flare · medopenai candidate 28 / 28 343 $0.0374 $37 1.2× 23 s 24 s 26 s 18 s 30 s 0 0 1,452 KB
sunburst · lowopenai candidate 28 / 28 158 $0.0316 $32 1.0× 26 s 26 s 29 s 19 s 31 s 0 0 1,420 KB
sunburst · medopenai candidate 28 / 28 343 $0.0372 $37 1.2× 27 s 28 s 31 s 21 s 32 s 0 0 1,448 KB
Time by category mean seconds per drawing type
CategoryDrawings2 · lowflare · lowflare · medsunburst · lowsunburst · med
Freehand scenes12524232423
Coloring page, magic brush13023222828
Coloring page, colored by hand12223252929
Coloring page, barely started12523262731
Crayon captures12222252628
Filled drawings12221202224
Outlines only12018202024
Magic brush on blank paper12622242627
Messy sessions11921202423
Night mode12623232529
Pretend-play probe12524243027
Scribbled fill12522222628
A few strokes, one color12525272528
Store scenes12222252731

Calls that returned no image

none

Every drawing came back as an image from every candidate. No refusals, no errors.

Gallery

every drawing, every candidate

Each drawing is followed by one candidate per column. The compact view shows the first trial; switch to both for more detail. Tap any result to swap it for the child's drawing and back, so you can see exactly what changed. Use the toolbar to hide candidates you have ruled out.

Trials
2
flare
sunburst

How this was measured