image-prompting
Author prompts for image-generation tools — covers, thumbnails, post visuals, illustrations, infographics, logos; intake→spec→prompt→verify loop.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Author prompts for image-generation tools — covers, thumbnails, post visuals, illustrations, infographics, logos; intake→spec→prompt→verify loop.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Schema and migration semantics for /dr-doctor — thin one-liner contract, 6-pass migration, data-loss safety, conflict resolution. Loaded by self-heal.
Core Datarim rules. Load this entry first, then only the fragment needed for paths, storage, numbering, backlog, routing, or archive behavior.
Post-QA hardening — detects task type (code, docs, research, legal, content, infra) and applies the matching verification checklist before archiving.
Testing pyramid, frameworks, mocking. Load first; then the fragment for the active gate (live smoke, silent failure, bats, legacy triage).
Preserve Datarim task continuity while orchestrated Claude Code or Codex sessions compact or clear context at deterministic pressure thresholds.
Immutability contract for all pipeline stages: artefact freeze, V-AC parity, non-code parity, anti-tautological rule, and return-to-source transition.
| name | image-prompting |
| description | Author prompts for image-generation tools — covers, thumbnails, post visuals, illustrations, infographics, logos; intake→spec→prompt→verify loop. |
| allowed-tools | Read Write Edit Grep Glob |
| model | inherit |
| current_aal | 1 |
| target_aal | 1 |
A reusable method for turning a content brief into a precise, repeatable prompt for a modern instruction-following image generator (the gpt-image family and equivalents). Load this whenever a task needs a visual asset: a blog cover, a video thumbnail, a social-post image, an inline illustration, a diagram/infographic, a logo mark, or an edit of an existing image.
This skill produces a prompt and a verification pass, not the image itself. It does not call any rendering API — it gives the caller the text to feed one, plus the size/quality settings to request and a checklist to judge the result.
A generator renders what it can resolve from the words it is given. Two failure modes dominate: under-specification (it invents the parts you left blank) and drift (across edits or iterations, the parts you wanted kept silently change). The whole playbook is built to defeat both:
intake → spec → prompt → render → verify → (refine one variable) → ship
Stay in spec until the brief is unambiguous. Most wasted renders trace to a thin spec, not a weak model.
Before writing a single prompt word, resolve these from the task brief (ask only if the answer changes the image; otherwise pick a sane default and note it):
| Slot | Question | Default if silent |
|---|---|---|
| Purpose | Cover, thumbnail, post image, illustration, infographic, logo, edit? | infer from where the asset will be used |
| Subject | What is literally in the frame? | the brief's headline noun |
| Placement | Where will it live (platform, page region)? | drives aspect ratio — see §10 |
| Text-in-image | Any words rendered inside the image? Exact string? | none — keep text out, overlay later |
| Brand/style anchor | A palette, a prior asset, a house style to match? | neutral, clean, modern |
| Mood | One or two adjectives for the feeling | matches the content tone |
| Hard noes | Anything that must NOT appear | no watermark, no stray text, no logos |
State the purpose explicitly in the prompt ("a blog cover image", "a UI mockup", "an infographic") — it switches the model into the right rendering mode and changes its defaults for layout, text density, and realism.
Order the prompt from the outside in, so the model establishes the scene before it places detail. A reliable slot order:
[purpose/medium] → [scene/background] → [subject + pose/action] →
[key details] → [composition/framing] → [light] → [palette/mood] →
[literal text, in quotes] → [negative constraints / invariants]
Format is flexible — a clean descriptive paragraph, a line-per-slot list, or a tagged/structured block all work. For anything you will reuse or hand to another agent, prefer the line-per-slot form: it is maintainable, diffable, and easy to vary one slot at a time. Cleverness in syntax buys nothing; clarity does.
A minimal prompt is often enough; add detail only where the default would be wrong. Long prompts work but are harder to debug — grow them by isolated additions, not by front-loading every adjective you can think of.
Composition is the cheapest lever on perceived quality. Specify:
For people and creatures, also pin body crop ("full body, feet included" vs "head and shoulders"), gaze ("looking off-frame, not at camera"), scale ("child-sized against the table"), and interaction ("hands gripping the handle"). These fix the three things generators get wrong most: proportion, action geometry, and where the eyes point.
State the visual medium first — it is the largest stylistic switch:
Then layer style qualifiers only as needed: brushstroke texture, film grain, halftone, cel shading, soft outlines, high detail vs minimal. Add quality levers (grain, macro detail, textured strokes) sparingly and only when the bare medium reads too clean.
Name a concrete reference register rather than an artist's name where possible ("editorial magazine illustration", "children's picture-book", "technical blueprint", "premium product photography") — it transfers a whole aesthetic without copying any one person and keeps the output reusable.
For character/brand consistency across a series: lock an anchor description once (appearance, proportions, outfit, palette, demeanor) and on every later image instruct "same character, do not redesign — new scene only", restating the anchor traits each time. Style continuity must be named, not assumed.
Reach for these only when the target is photographic realism:
Light sets realism and mood more than any other single factor:
For edits that change weather or time of day, change only the light and atmosphere and explicitly preserve camera angle, object positions, and scene identity.
Rendering legible text is the hardest thing these models do. When the brief truly needs words baked into the image:
the word "LAUNCH" in bold sans-serif.Default to keeping text OUT of the generated image and compositing it in a layout tool afterward. It is more legible, more on-brand, and trivially editable. Bake text in only when it must interact with the scene (a sign, packaging, a poster within the world).
Two distinct tools — do not conflate them:
Rules for invariants:
Modern instruction generators render at a native size you request, rather than a fixed square you upscale. Request the aspect ratio the placement actually needs and let the tool render it natively:
| Placement | Aspect | Typical native size |
|---|---|---|
| Blog/article cover, OG image | 16:9 / 1.91:1 | landscape ~1536×1024 |
| Video thumbnail | 16:9 | landscape ~1536×864–1024 |
| Vertical story / reel / pin | 9:16 / 2:3 | portrait ~1024×1536 |
| Square post | 1:1 | 1024×1024 |
| Slide / deck | 16:9 | landscape ~1536×864 |
Practical constraints for the gpt-image family (verify against your tool's current limits before shipping a pipeline):
Quality vs. cost tiers:
Edit prompts are a special case of §9: the scene already exists, so almost everything is an invariant. The shape is always:
[the single change] + [keep everything else identical] + [enumerate the untouchables]
Common edits and their one-line forms:
When the tool supports it, request high input fidelity for edits that must hold identity through a big change.
Fill-in-the-blank prompt skeletons for the common asset types (cover, thumbnail, social post, illustration, infographic, logo, photoreal, edit) live in prompt-templates.md. Copy the one matching the purpose, fill the slots from the §1 intake, and tune one variable at a time.
Run before shipping any generated asset:
Related skills, load when relevant:
writing / publishing — when the image accompanies an article or post; align mood and palette with the copy and respect platform image specs.frontend-ui — when the asset is a UI mockup or must match a site's theme tokens and aspect ratios.brainstorming — when the visual concept itself is open and needs exploration before a spec exists.