| name | figure |
| description | Turn article passages, technical concepts, or short explanations into 16:9 white-background, mostly black-and-white, hand-drawn line illustrations using Codex image generation/image-2. Use when the user wants a passage, section, or paragraph in Chinese, English, or mixed language transformed into visual explanation shots, methodology diagrams, process/state/action illustrations, or a shot list where a default hand-drawn rabbit IP participates in the core cognitive action rather than decorating the scene. |
Figure
Overview
Use this skill to extract the single most important cognitive structure from article passages, technical concepts, or short explanations and turn it into focused explanatory illustration shots. The default output is a 5-8 shot list for full articles, or one generated 16:9 image for a single paragraph/passage.
Core Rules
- Treat the illustration as a focused thinking diagram, not an article thumbnail or instruction manual.
- First converge on the single most important core viewpoint. Do not try to include every detail, example, tool, role, or branch from the source text.
- Each image must explain one core structure only: one flow, one state change, one comparison, one method, one system relation, or one decision.
- Identify the cognitive action after convergence: compare, sort, filter, connect, iterate, diagnose, choose, compress, expand, map, hand off, break down, or reframe.
- Make the rabbit IP perform the core action. Do not place the rabbit as a mascot, narrator, sticker, corner decoration, or unrelated observer.
- The rabbit's gender presentation is flexible: infer a suitable gender expression from the passage when the content clearly implies one; otherwise keep the rabbit gender-neutral.
- Preserve the rabbit's character proportions in every final prompt using the Character Consistency Lock from references/visual-ip.md: a small hand-drawn rabbit with the same loose editorial proportions as the user's preferred reference style — a soft oval head only slightly larger than the torso, not an oversized toddler head; a small simple shirt-shaped torso, short pants or a tiny vest when useful, short thin arms, short thin legs, and small rounded hands and feet. The body stays compact, light, and a little wobbly, with cuteness coming from loose hand-drawn posture and clothing details rather than a huge head or chubby body.
- Keep the rabbit visually distinct from Miffy / Nijntje and similar classic minimalist picture-book rabbits: avoid the exact combination of symmetrical upright ears, centered dot eyes, and an x-shaped mouth. Use natural uneven ears, a tiny off-center short horizontal dash mouth, a small subtle nose, and small clothing details to make the character feel like original editorial line art rather than an existing character.
- The rabbit mouth is a hard constraint: it must be exactly one tiny short horizontal dash, drawn as a single left-to-right pen stroke. Never draw a cross, X, plus sign, two crossing strokes, V shape, smile curve, frown curve, dot mouth, open mouth, or any multi-stroke mouth mark.
- Use a pure white background, mostly black-and-white line art, hand-drawn marker/pen texture, and a 16:9 wide composition.
- Color is allowed only as small muted Korean-illustration-style clothing color blocks on the rabbit. Keep all diagrams, labels, arrows, backgrounds, and information structures black-and-white. Leave a thin white gap between each clothing color block and the black outline.
- Keep text labels sparse, short, legible, and embedded in the diagram. Prefer 2-5 labels per image.
- Match label language to the source text by default: Traditional Chinese for Traditional Chinese input, Simplified Chinese for Simplified Chinese input, concise English for English input, and preserve key technical terms in their original language for mixed-language input.
- Render text labels directly in the image by default. The handwritten style controls only how the strokes look — every character must still be written correctly and stay clearly legible, with no missing or extra strokes, no garbled glyphs, and no look-alike or near-homophone substitutions. Note that image-2 renders Chinese glyphs imperfectly; keep Chinese labels short (favor 2-3 over 4-5) to reduce errors.
- Use this label style exactly: 中文標註風格:字形連貫流動,筆畫帶速度感,線條有自然抖動和停頓,可能會因為急促感而將幾個筆畫連起來,但不影響閱讀,筆尾輕輕拉長但不過於誇張,結構鬆動但清晰,整體像情緒很滿時寫下的一句手寫標題。
- This skill does not bundle a font. Post-processing labels with scripts/render_labels.py is an advanced, opt-in path only for when the user explicitly asks for stable typography AND supplies their own handwriting
.ttf (passed as the script's 4th argument).
- Avoid color outside the rabbit's small muted clothing color blocks unless the user explicitly overrides the default.
- Prefer Traditional Chinese labels when the input uses Traditional Chinese; mirror the input's language and Chinese variant when clear.
Workflow
- Classify the input:
- One paragraph or a selected passage: create 1 shot spec and generate 1 illustration unless the user only asks for a prompt or shot list.
- A full article or multiple sections: create a 5-8 shot list with prompts; generate all images only if the user explicitly asks for image generation for the full set.
- If uncertain, default to 1 shot for short inputs under roughly 500 Chinese characters or 300 English words, and 5-8 shots for longer structured articles.
- Extract visual candidates:
- Find claims that describe movement, transformation, tension, workflow, state change, decision-making, or method.
- Skip purely atmospheric, anecdotal, or decorative lines unless they reveal the article's core thinking.
- Rank candidates by explanatory value, then keep only the most important one for each image.
- Choose each shot's structure type:
flow: steps, sequence, pipeline, feedback loop.
state-map: before/after, current/future, stuck/unstuck, health/status.
comparison: A/B, tradeoffs, spectrum, matrix.
method: framework, checklist, layered model, operating principle.
system: roles, inputs/outputs, dependencies, handoff.
decision: fork, criteria, prioritization, constraint sorting.
- Design the rabbit's action:
- Make the rabbit physically manipulate the idea with compact, stable body mechanics: tap, stamp, place, slide, circle, pin, underline, hold a card close to the chest, press a button, nudge a block, point with a short pointer, or connect two nearby dots with a short line.
- Keep every hand action inside the rabbit's short-arm reach. The elbow should stay visibly bent or close to the torso; the writing hand must not extend farther than about one head-width away from the body.
- Prefer poses with both feet planted, sitting, kneeling, leaning lightly against a prop, or standing on a small stool. Avoid poses that require long reaching, stretching across the page, pulling distant strings, climbing, twisting, balancing on one foot, or holding a tool far from the body.
- For writing, highlighting, drawing arrows, or connecting lines, place the rabbit close to the target mark and move the page/card/diagram element near the rabbit rather than extending the rabbit's arm across the composition.
- Ensure the action changes or reveals the information structure.
- Give the rabbit a creative editorial gesture only when it remains anatomically simple, grounded, and readable. The pose should feel intentional, not contorted.
References