| name | write-image-prompts |
| description | Use when prompting an image-generation model: a hero image, icon set, product shot, illustration, OG image, scene, or any generated visual. Covers concrete specifics over quality adjectives, specifying the world not just the subject, describing presence not absence, reference images over style prose, anchoring sets, iterating by editing, and designing for the image's job. The image sibling of write-prompts. |
Write image prompts: describe what the camera sees, not how good it should be
Image models render nouns, parameters, and relationships. They do not render aspirations. Most weak image prompts fail the same way weak text prompts do — theatre instead of information — but the fixes are image-specific.
Concrete specifics beat quality adjectives. "Professional, stunning, high quality, 8k masterpiece" are vibe words; the model already tries its best. What actually steers: photography vocabulary ("85mm f/1.8, shallow depth of field, golden-hour backlight, eye level"), material and texture words, named composition ("rule of thirds, subject left, negative space right"). These are the levers the training data encodes. Same per-line test as write-prompts: does this word change the pixels, or just the vibe?
Specify the world, not just the subject — and know that labels summon stereotypes. Every attribute you don't state regresses to that model's default world, and the default leans hard American: probed mid-2026, a frontier model given just "a suburban house on a quiet street" and "a small town main street" returned US flags in both images, unprompted, plus American road markings and an American drugstore. The lean varies by model and generation, so find yours the same way — generate a deliberately underspecified scene and look at what arrives. And a country or culture label doesn't fix an unstated world, it summons the postcard: "Australian" gets red dirt, tin roofs, the Harbour Bridge — not your actual suburb — and the same cliché-retrieval applies to any place or culture. Say the real details concretely instead: the brick veneer, the light, the brands of equipment, which direction the sun comes from. This is silent failure either way — subtly wrong everywhere, or loudly stereotyped in a way locals clock instantly.
Specification is a dial, because image models are literal. Whatever you name will appear, and prominently — say "gum trees" and gum trees you shall have, front and centre. So unstated regresses to stereotype, but stated gets over-rendered. The craft is choosing the few elements that must be right, stating those, and letting a reference image carry the rest of the world rather than enumerating it.
Describe presence, not absence. Negative instructions are weak — "no text", "no people", "don't make it cluttered" often fail. Describe the composition you want so fully that the unwanted thing has nowhere to be: "a clean empty benchtop, single product centred" beats "no clutter". On text-capable models, also keep instructions and content cleanly separated: stray instruction words can get rendered into the image.
A reference image beats a paragraph of style prose — and it's the non-photographer's lever. One real image anchors palette, lighting, and rendering style better than any description — the worked-example rule in image form. If you don't know the photography vocabulary, you don't need it: the reference image is the vocabulary. This is also how sets and recurring subjects stay consistent: generate the first piece with full direction, then pass it as the reference for the rest. Asking for N variations gives you N takes on one prompt, not N different subjects; and generating a sheet of items to slice up fails on alignment — generate each item isolated, anchored to the shared reference. For anything recurring (a brand, a mascot, a location, a product line), build a small library of your own grounding shots and characters once — world-building is the highest-leverage asset in image generation, because every future prompt gets to point at it instead of describing it.
Iterate by editing, not re-rolling. When a result is close, multi-turn editing from the current image ("same scene, swap the jacket to navy") preserves what's working; regenerating from a tweaked prompt re-rolls everything. Change one thing per turn so you can tell what each change did. And iterate with the image, not by asking another model to talk the prompt better: refinement-by-discussion produces procedural jargon ("zero edge contrast", "interlock at the cut") that image models, trained on captions of what a scene is, render as scene content (write-prompts covers when interviewing a model does help).
Ground real places in real images, not words. Ask for "Newcastle Harbour, Australia" by words alone and you get the place as someone might paint it from a dream of having been there — plausible, wrong in every particular. The fix, in order of strength: your own reference photos of the actual place (hard to beat), a model's image/search grounding where offered, and at minimum researching the subject and feeding observed concrete details into the prompt. The truth-seeking rule applies to pixels too: a real place is a checkable claim.
Design for the image's job, not just its content. An image has a slot: a hero needs copy space where the headline goes and must survive a crop; an icon needs to read at 24px, single subject, flat; an OG image needs the focal point centre-safe. Specify aspect ratio, where text will overlay, and what it must look like small. A beautiful image that fails its slot is a failed image.
Models differ wildly per job — verify live, then look. Text rendering, transparency support, faces, speed, and editability vary by model and shift monthly; pick by bake-off on your actual task, not reputation (verify-current). And every generated image gets looked at before it ships — text-in-image especially needs a QA read, because a small percentage of renders carry typos that spell-check will never see (verify-visually).
The failure this prevents: a folder of generic stock-looking output from a capable model, a "local business" image set that's subtly the wrong country, an icon set where every icon has a different line weight, a hero with no room for the headline, and a typo shipped inside a PNG. Pairs with write-prompts (the same theatre test, for text), reason-over-images (the model reading images; this is the model making them), and verify-visually (nothing visual ships unlooked-at).