- name
- academic-figure-generator
- description
- Create publication-ready computer science figures with gpt-image-2 through a user-configured OpenAI-compatible image endpoint. Use when the user asks to draw, illustrate, visualize, or generate a paper figure such as a system architecture, workflow, technical mechanism, algorithm intuition, taxonomy, related-work comparison, ablation, or benchmark overview. The skill first reconstructs the figure's factual argument and layout, then writes a structured image prompt, calls the configured endpoint, inspects the result, and retries with targeted corrections when the image is not faithful or publication-ready.
# Academic Figure Generator
Generate figures as arguments, not decorations. The figure must make one technical claim legible within a few seconds and must preserve the user's facts, labels, directions, and comparisons.
## Operating contract
Run this sequence for every generation request:
1. **Extract the figure thesis.** State in one sentence what the reader should learn from the figure. Select exactly one primary figure type: `architecture`, `pipeline`, `mechanism`, `comparison`, `taxonomy`, `timeline`, `ablation`, or `overview`.
2. **Build a fact ledger.** Separate user-provided facts from agent inferences and visual choices. Preserve every exact name, number, ordering, dependency, baseline, metric, and relationship in the fact ledger. Uncertain facts stay marked as unknown and are not silently completed.
3. **Resolve blocking ambiguity.** Ask only for information that changes the graph or the claim: missing input/output, unclear arrow direction, undefined comparison axis, contradictory labels, or an unknown method name. For style-only gaps, choose a restrained default and continue.
4. **Create a layout specification.** Decide the reading direction, regions, hierarchy, visual encoding, label budget, and output aspect ratio before writing the prompt. The layout must have one dominant path and a clear visual start and end.
5. **Write the generation prompt.** Use the prompt template below. Put the factual specification before style language. Use concrete geometry and spatial relations instead of vague words such as `beautiful`, `cool`, or `professional`.
6. **Save the prompt artifact.** Write the final prompt to `<output-stem>.prompt.txt` and, when useful, write the fact ledger and layout specification to `<output-stem>.spec.json` before calling the image endpoint.
7. **Generate the image.** Run `scripts/generate_image.py` with the prompt file and an explicit output path. Read the endpoint and key from the private user configuration created by `scripts/configure.py`; never put credentials in the skill directory.
8. **Inspect the result.** Open the generated image with the image inspection capability available to the agent. Check semantic correctness, text fidelity, arrow direction, occlusion, hierarchy, whitespace, contrast, and print-scale legibility. Compare against the fact ledger, not against the model's interpretation.
9. **Repair or reject.** If a critical fact, label, relationship, or direction is wrong, rewrite the smallest relevant prompt section and regenerate. Do not rationalize a wrong diagram. If the model repeatedly corrupts exact text, report that limitation and recommend a vector/SVG or slide-based finalization for the labels.
10. **Report reproducibility data.** Return the image path, prompt path, model, non-secret configuration path, size, quality, number of attempts, and any unresolved accuracy risk. Never print or include the API key.
A run is complete only when the image exists, opens as a valid image, matches the fact ledger, and passes the publication checklist below. A prompt without a generated image is complete only in explicit `brief-only` mode.
## Figure modes
Use the smallest mode that answers the request:
- `brief-only`: produce the thesis, fact ledger, layout specification, and final prompt without calling the API.
- `generate`: produce the prompt, call the API, inspect the image, and repair critical defects.
- `revise`: use the user's existing figure or critique as the source of truth, preserve accepted content, and change only named defects. If image-edit support is unavailable at the configured endpoint, create a fresh prompt that explicitly preserves the accepted structure.
- `variant`: generate controlled alternatives that differ in one dimension only, such as reading direction, grouping, or visual encoding. Keep the fact ledger identical across variants.
## Fact ledger
Construct this internally before generation, and expose it in the prompt artifact when the user asks for an auditable workflow:
```json
{
"thesis": "one falsifiable reader takeaway",
"figure_type": "architecture|pipeline|mechanism|comparison|taxonomy|timeline|ablation|overview",
"audience": "systems or ML researchers",
"entities": [
{"id": "e1", "label": "exact visible label", "role": "input|module|state|output|baseline|metric"}
],
"relations": [
{"from": "e1", "to": "e2", "kind": "data|control|dependency|comparison|feedback", "label": "optional exact text"}
],
"ordered_steps": [],
"comparisons": [],
"numbers_and_units": [],
"must_preserve": [],
"uncertainties": [],
"layout": {
"direction": "left-to-right|top-to-bottom|radial|matrix",
"regions": [],
"aspect_ratio": "16:9",
"target_use": "two-column paper|single-column paper|slide"
}
}
```
Use exact labels from the ledger. Keep labels short by moving explanations into callouts, but do not abbreviate a technical term unless the user supplied the abbreviation. Do not invent baselines, metrics, datasets, module names, causal links, performance values, or citations.
For related-work figures, define the comparison axes before styling. Each method must have the same fields, the axes must be named, and an empty or unknown value must be rendered as `N/A` or omitted according to the user's instruction. A visual position on an axis is not evidence of superiority unless the source facts support it.
## Layout rules by figure type
### Architecture and pipeline
Use a single dominant reading path. Place inputs on the left or top, transformations in the center, and outputs on the right or bottom. Group modules into 3 to 5 labeled stages. Put feedback, side information, and control paths on distinct routed lines. Make data flow and control flow visually distinct, and give every crossing a visible bridge, gap, or explicit junction.
### Technical mechanism
Show the before-state or bottleneck, the proposed intervention, and the resulting state. Use one zoomed inset only when it explains the mechanism rather than adding ornament. Encode time, memory, locality, or hierarchy with an explicit legend when those dimensions matter.
### Related work and comparison
Use a matrix, aligned lanes, or a small coordinate chart with a legend. Give every method equal visual area and equal label treatment. Put the proposed method in a restrained accent, and keep the comparison claim textual and qualified by task, data, metric, and budget.
### Taxonomy, timeline, ablation, and overview
Use consistent levels, aligned branches, or matched panels. Keep each branch mutually interpretable. For ablations, make the changed component and evaluation quantity explicit; do not turn a numerical result into a decorative badge. For overviews, preserve a clear hierarchy from problem to method to evidence.
## Prompt construction
Write the prompt in this order:
1. `Task and thesis`: the exact reader takeaway and figure type.
2. `Canvas`: aspect ratio, white or transparent background, margin, print-safe contrast, and no crop.
3. `Composition`: reading direction, regions, alignment, grouping, and whitespace.
4. `Objects`: one entry per entity with shape, position, role, and exact label.
5. `Relations`: one entry per edge with source, destination, arrow direction, line type, and exact edge label.
6. `Visual encoding`: semantic color roles, stroke hierarchy, line styles, and legend.
7. `Text contract`: exact spelling, readable Latin text, consistent typography, no invented text, no duplicated labels, and no text placed over lines or shapes.
8. `Output constraints`: clean vector-like scientific illustration, publication-ready, no perspective, no photorealism, no decorative background, and no watermark.
Use this scaffold, replacing every bracketed field:
```text
Create a publication-ready computer science paper figure.
THESIS:
[one sentence describing the reader takeaway]
FIGURE TYPE:
[one type]
AUDIENCE AND USE:
[venue/audience], [single-column, two-column, or slide], [aspect ratio]
CANVAS:
A flat, clean, high-resolution scientific diagram on a white or transparent background. Use generous outer margins and a balanced composition that remains legible when reduced to [target width]. Use a left-to-right [or chosen] reading direction.
COMPOSITION:
[describe the regions and the dominant path in spatial order]
[describe grouping, alignment, whitespace, inset, legend, and any deliberate crossing treatment]
ENTITIES, PRESERVE THESE EXACT LABELS:
1. [label]: [role, shape, location, and visual meaning]
2. [label]: [role, shape, location, and visual meaning]
...
RELATIONS, PRESERVE THESE EXACT DIRECTIONS:
- [source] -> [destination]: [data/control/dependency/comparison], [solid/dashed/double line], [exact edge label or no label]
...
FACTS:
- [exact number, unit, ordering, baseline, metric, or condition]
...
VISUAL ENCODING:
Use a restrained academic palette with distinct semantic roles: [role/color]. Use color as a secondary encoding with a grayscale-readable contrast. Use one accent color for the focal contribution and neutral colors for context. Include a compact legend only if the encoding cannot be inferred locally.
TEXT CONTRACT:
Render every listed label exactly once with exact spelling and capitalization. Use readable typeset Latin text, consistent sans-serif typography, short labels, aligned baselines, and clear separation from arrows and borders. Keep all labels inside their containers or in nearby callouts. Use only the listed text and the required legend text.
OUTPUT:
Crisp vector-like geometry, consistent stroke widths, aligned edges, uncluttered whitespace, publication-ready scientific communication. No decorative objects, visual effects, gradients, shadows, pseudo-3D perspective, photorealistic imagery, watermark, logo, citation, or unrequested text.
```
The model prompt is a rendering instruction, not a place to make unsupported claims. Every visual emphasis must have a corresponding role in the fact ledger. Use pale fills, dark text, 1.5 to 2.5 px equivalent strokes, and one accent color by default. Adjust colors for color-vision accessibility and grayscale reproduction.
## Accuracy and inspection gate
Treat these as severity levels:
- **Blocker:** wrong or missing entity, wrong arrow direction, false causal relation, fabricated number, incorrect method name, unreadable essential label, cropped content, or a comparison that visually asserts an unsupported conclusion.
- **Major:** ambiguous grouping, misleading color semantics, overlapping text, unclear start/end, inconsistent panels, or a legend that changes interpretation.
- **Minor:** spacing imbalance, redundant ornament, weak alignment, or typography that is acceptable but not polished.
A blocker requires regeneration. Two majors in the same area require regeneration. Minor issues may be reported when the image is otherwise usable.
Inspect at both native resolution and the intended paper width. Run this checklist:
- The thesis is visually dominant and understandable without reading a paragraph.
- Every ledger entity appears exactly once unless repetition is explicitly specified.
- Every required relation has the correct direction and line semantics.
- Inputs, intermediate states, outputs, and feedback are distinguishable.
- No label, arrowhead, border, or legend is clipped or occluded.
- Exact labels, numbers, units, and capitalization match the ledger.
- Color carries meaning consistently and remains interpretable in grayscale.
- Alignment, spacing, stroke weight, and panel treatment are consistent.
- The figure has no unsupported visual ranking or invented evidence.
- The image is a valid PNG, JPEG, or WebP and has enough resolution for its target use.
When exact text fails after two targeted retries, stop spending attempts on prose variations. Keep the generated visual structure as a draft and recommend finalizing text, arrows, and geometry in SVG, Illustrator, Figma, Inkscape, PowerPoint, or LaTeX/TikZ. Accuracy outranks the appearance of completion.
## API execution
Use the bundled scripts rather than writing ad hoc HTTP code. Configure the relay once with hidden key input:
```bash
python3 scripts/configure.py
```
The configuration is stored outside the skill and source repositories:
- Linux and macOS: `${XDG_CONFIG_HOME:-~/.config}/academic-figure-generator/config.json`
- Windows: `%APPDATA%\\academic-figure-generator\\config.json`
On POSIX systems the file is written with mode `0600`, and generation refuses to read an API key from a more permissive file. The JSON supports `api_url`, `api_key`, `model`, `size`, `quality`, `background`, `output_format`, `timeout`, and `retries`. `api_url` may be a full images generation endpoint or an OpenAI-compatible `/v1` base URL.
Generate with the saved configuration:
```bash
python3 scripts/generate_image.py \
--prompt-file figure.prompt.txt \
--output figures/system-architecture.png
```
Use `--config <path>` for a separate profile and explicit command-line options to override non-secret settings for one run. Environment variables remain a backward-compatible fallback but are not the recommended configuration method. The script accepts `b64_json`, a returned image URL, or a direct image response. The key is sent as a bearer token and is never included in prompt artifacts or normal logs.
## Default deliverables
For `generate`, provide:
- the final image path;
- the exact prompt artifact path;
- a concise figure specification with thesis, figure type, and output settings;
- the inspection result, including any remaining blocker, major, or minor issue;
- a note that exact text and semantics were checked against the fact ledger.
For `brief-only`, provide the complete fact ledger, layout specification, and prompt, and explicitly state that no API call was made.
Voir sur GitHub