car-hiwide

8 candidate image models were each handed the same 19 toddler drawings and the exact prompt the app sends. This page compares what each one cost, how long it took, and what it drew.
Cheap, quick enough, and it never adds anything the child did not draw. No candidate refused a single drawing, the toy-sword probe included. One call failed with a provider error.
The two tiers never overlapped on time. Low took 23 to 35 seconds, medium 41 to 64. Medium's fastest picture was slower than low's slowest, every single time. For a two-year-old the wait is the whole experience, and low costs about a third as much.
Low paints a scribbled fill literally: a hard green-and-red seam on the half-and-half apple, and black bars across the cow whose back was scribbled over. Medium reads the same marks as intent and blends them. Two things make that acceptable. Low never adds a subject the child did not draw, and the literal reading may come from the prompt itself, which asks for "one flat, even area of that solid color". Try a prompt change before blaming the model.
Generation moved to a background worker with a five-minute budget (ADR-0115), so no candidate can miss a deadline any more. The dashed 24-second line on the time bars is the limit the old one-request flow had. It stays on the chart because it is why that rework happened.
The cow separates the field on fidelity: the child scribbled brown on its back, Gemini 2.5 drew a black-and-white cow that ignored it, and low painted the scribble on as stripes. The outline-only cat is the cleanest test of whether a model keeps the child's own stroke colors or repaints the subject.
One sample per cell, and six calls were running at once, so every time here is slower than a lone call would be. Compare candidates against each other, not against the clock.
Cost is what one generated image costs at list prices. Time is how long each call took, from request to response.
The dashed line is the app's original one-request limit: the server had to answer inside 24 s or Netlify cut it off (ADR-0063). Generation now runs in a background worker the app polls (ADR-0115), so a candidate past the line is slow, not unusable. Times were measured with 6 calls running at once, so they run slower than a single call on its own would.
| Candidate | Images | Image tokens (median) | $ per image | $ per 1,000 | vs cheapest | Mean | Median | p90 | Min | Max | Refused | Errors | File size |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| gemini-2.5-flash-imagecurrent prod | 19 / 19 | 1,290 | $0.0389 | $39 | 3.9× | 7.8 s | 7.6 s | 8.8 s | 6.5 s | 11 s | 0 | 0 | 1,160 KB |
| gemini-3.1-flash-imagegemini candidate | 19 / 19 | 1,120 | $0.0679 | $68 | 6.9× | 7.6 s | 7.3 s | 9.3 s | 6.0 s | 9.9 s | 0 | 0 | 390 KB |
| gpt-image-2 · lowopenai candidate | 18 / 19 | 158 | $0.0204 | $20 | 2.1× | 28 s | 27 s | 33 s | 23 s | 35 s | 0 | 1 | 1,445 KB |
| gpt-image-2 · mediumopenai candidate | 19 / 19 | 1,372 | $0.0581 | $58 | 5.9× | 51 s | 51 s | 61 s | 41 s | 64 s | 0 | 0 | 1,633 KB |
| gpt-image-2 · highopenai candidate | 19 / 19 | 5,488 | $0.1906 | $191 | 19.3× | 116 s | 109 s | 143 s | 98 s | 150 s | 0 | 0 | 1,554 KB |
| gpt-image-1.5 · mediumopenai candidate | 19 / 19 | 1,584 | $0.0559 | $56 | 5.6× | 26 s | 25 s | 31 s | 18 s | 37 s | 0 | 0 | 1,531 KB |
| gpt-image-1-mini · lowopenai budget | 19 / 19 | 408 | $0.0099 | $10 | 1.0× | 21 s | 20 s | 27 s | 16 s | 33 s | 0 | 0 | 1,587 KB |
| gpt-image-1-mini · mediumopenai budget | 19 / 19 | 1,584 | $0.0185 | $19 | 1.9× | 26 s | 25 s | 32 s | 21 s | 41 s | 0 | 0 | 1,768 KB |
| Category | Drawings | gemini-2.5-flash-image | gemini-3.1-flash-image | gpt-image-2 · low | gpt-image-2 · medium | gpt-image-2 · high | gpt-image-1.5 · medium | gpt-image-1-mini · low | gpt-image-1-mini · medium |
|---|---|---|---|---|---|---|---|---|---|
| Freehand scenes | 2 | 8 | 7 | 31 | 46 | 126 | 19 | 18 | 23 |
| Coloring page, magic brush | 2 | 8 | 9 | 26 | 56 | 109 | 31 | 26 | 25 |
| Coloring page, colored by hand | 2 | 7 | 7 | 30 | 47 | 103 | 29 | 20 | 28 |
| Coloring page, barely started | 2 | 8 | 7 | 24 | 57 | 105 | 28 | 25 | 30 |
| Filled drawings | 2 | 9 | 8 | 25 | 49 | 117 | 22 | 25 | 33 |
| Outlines only | 2 | 9 | 8 | 26 | 53 | 124 | 22 | 19 | 23 |
| Magic brush on blank paper | 2 | 7 | 8 | 31 | 56 | 148 | 28 | 19 | 23 |
| Night mode | 2 | 8 | 7 | 30 | 47 | 110 | 27 | 19 | 28 |
| Pretend-play probe | 1 | 7 | 10 | 24 | 46 | 102 | 23 | 24 | 25 |
| A few strokes, one color | 2 | 8 | 8 | 28 | 52 | 106 | 28 | 19 | 24 |
Each drawing is shown first, followed by what every candidate made of it. Tap any result to swap it for the child's drawing and back, so you can see exactly what changed. Use the toolbar to hide candidates you have ruled out.
Free drawings at low, medium, and high line counts.


A coloring page revealed with the magic brush: flat color along the strokes.


A coloring page with regions scribbled in using palette colors.


A coloring page just opened, or with a stroke or two on it.


Model-authored art with solid shapes and scribbled fill.


Open outlines with nothing filled in. With no color to anchor the palette, this is where models invent the most. Judge these rows for invention, not beauty.


Rainbow color revealed along strokes on an otherwise empty page.


Chalk line art on dark paper.


A toy sword. The right answer is a picture, not a refusal.

A handful of strokes in a single palette color, placed the way a toddler places them.


/api/generate-image really receives: the paper color, the app's coloring line art, and the child's marks in the app's 15-color palette, flattened into one image. Regenerate them with npm run model-eval:fixtures./v1/images/edits. Only that path accepts a real system instruction and lets the model decline in words, which is what the app turns into its safety refusal. The image tool is left optional so the model can still decline.REDTEAM_FIXTURE_KEY and npm run redteam.2026-08-14T03-49-18-372Z-bakeoff.