Skip to main content

ai-image-prompt-debugging

Diagnoses and fixes AI image generation prompt issues including mismatched style descriptors, conflicting terms, underspecified composition, and overcrowded token budgets with revised prompt output. Use when the user asks why their AI image prompt is not producing expected results, wants to fix a failing prompt, or needs to diagnose visual artifacts in generated images. Do NOT use for creating new prompts from scratch (use model-specific prompting skills), style transfer (use ai-image-style-transfer), or upscaling (use ai-image-upscaling).

Informations de source

Dépôt
FerroxLabs/murage
Dernière activité de la source
1 septembre 2026 à 13:26
Langue détectée de SKILL.md
anglais
Étoiles
9
Forks
2

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Explorateur de fichiers
2 fichiers

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
ai-image-prompt-debugging
description
Diagnoses and fixes AI image generation prompt issues including mismatched style descriptors, conflicting terms, underspecified composition, and overcrowded token budgets with revised prompt output. Use when the user asks why their AI image prompt is not producing expected results, wants to fix a failing prompt, or needs to diagnose visual artifacts in generated images. Do NOT use for creating new prompts from scratch (use model-specific prompting skills), style transfer (use ai-image-style-transfer), or upscaling (use ai-image-upscaling).
license
Apache-2.0
metadata
{"author":"foundry-skills","version":"1.0.0","tags":"ai-image-generation analysis design","category":"design-creative","subcategory":"ai-image-generation","depends":"","disclaimer":"none","difficulty":"intermediate"}
# AI Image Prompt Debugging ## When to Use **Use this skill when:** - A user provides an existing prompt and describes a gap between what they expected and what the model produced (wrong style, wrong subject, wrong composition, wrong mood, artifacts, missing elements) - A user asks why their prompt "used to work" in an older model version but produces degraded results after a model update - A user describes specific visual failure patterns: anatomy errors, style bleed, background dominance, color casts, texture corruption, or compositional collapse - A user has tried multiple prompt variations and cannot isolate what is causing the persistent problem - A user wants to understand the mechanics of why two nearly identical prompts produce drastically different outputs - A user's negative prompt is actively suppressing desired elements or failing to exclude unwanted ones - A user reports that CFG, step count, or sampler changes are producing unexpected responses to the same prompt - A user's weighted terms are causing visible artifacts, halos, or oversaturated focal points **Do NOT use when:** - The user has no prompt yet and wants to build one from scratch -- use a model-specific prompting skill instead - The user wants to migrate a working Midjourney prompt to Stable Diffusion or vice versa -- use a prompt-translation skill - The user wants to apply a specific visual style to a new subject they have not yet prompted -- use ai-image-style-transfer - The user wants to increase resolution or recover detail in an already-generated image -- use ai-image-upscaling - The user wants to generate variations of a working prompt -- use ai-image-variation or model-specific iteration skills - The user is troubleshooting inpainting or outpainting failures -- use ai-image-inpainting-debug - The problem is a model crash, API error, NSFW filter block, or queue failure -- this is infrastructure, not prompt diagnosis --- ## Process ### Step 1 -- Gather All Debugging Inputs Before diagnosing, collect every variable that affects output. Missing even one can lead to incorrect diagnosis. - **Exact prompt text:** Ask for a copy-paste, not a paraphrase. A single word difference can change output significantly. If the user is on Midjourney, ask them to use the `/describe` output or copy directly from the job history. - **Negative prompt (if SD/ComfyUI/InvokeAI):** The full negative prompt. A negative prompt that contradicts the positive is one of the most common causes of failure in Stable Diffusion pipelines. - **Model and version:** SD 1.5 vs SDXL 1.0 vs SDXL Turbo vs SD 3 vs Midjourney v5/v6/v6.1 vs DALL-E 3 vs Flux.1 [dev] vs Flux.1 [schnell] -- these are architecturally different and share almost no prompt syntax conventions. - **Checkpoint or LoRA (SD):** The base checkpoint (e.g., Realistic Vision v6, DreamShaper XL, Juggernaut XL) affects how prompt terms resolve to visual output. A prompt tuned for Deliberate v3 will fail on SDXL without adaptation. - **Sampler and scheduler:** DPM++ 2M Karras, Euler a, DDIM, UniPC, LCM -- each has different convergence behavior and responds differently to CFG values. - **CFG scale / guidance scale:** The exact number. This is diagnostic. - **Step count:** The exact number. Under 20 steps with most non-distilled samplers produces unfinished output regardless of prompt quality. - **Resolution:** Width x height. SDXL artifacts above 1.5 megapixels without MultiDiffusion or tiled VAE. SD 1.5 optimal is 512x512 or 512x768. Midjourney --ar ratio. - **Seed (if known):** A fixed seed lets you isolate prompt changes from randomness. - **What the user expected vs. what they got:** Ask for a description of the actual output, or better, the image itself. "It looks cartoonish" is less useful than "it rendered as a flat 2D illustration with cel shading and dark outlines." - **What variations they already tried:** This prevents recommending things they have already ruled out. --- ### Step 2 -- Identify the Model's Prompt Paradigm Different models process prompts through fundamentally different architectures. Diagnosis must start by confirming which paradigm governs this prompt. **Stable Diffusion 1.x (CLIP ViT-L/14 encoder):** - Hard token limit of 77 tokens per pass. Tokens are not words -- a word like "photorealistic" may be 3 tokens. Punctuation costs tokens. - Terms beyond token 77 are passed to a second CLIP pass with reduced influence. Many implementations simply truncate. - Comma-separated keyword style works best. Natural language sentences are processed less reliably. - Supports attention weighting: `(term:1.2)` increases emphasis, `(term:0.8)` reduces it. Values above 1.5 cause visual saturation/artifacts. Values below 0.5 cause near-suppression. - Checkpoint choice dominates style more than prompt terms. A LoRA with 0.8 weight can overpower a prompt entirely. - Highly sensitive to quality booster tags: "masterpiece, best quality, ultra-detailed" strongly bias SD 1.5 models trained on Danbooru/e621 datasets. **Stable Diffusion XL (two CLIP encoders: ViT-L and ViT-bigG):** - Two separate 77-token passes: one to each CLIP encoder. The ViT-bigG encoder (second) carries more semantic weight. - Responds better to natural language phrases within keyword strings. - Quality booster tags ("masterpiece") have less effect on SDXL than on SD 1.5. - Optimal base resolution: 1024x1024. Using 512x512 causes composition collapse and face malformation. - SDXL Refiner pipeline expects a denoising_start cutoff between 0.7-0.9 (approximately the last 20-30% of steps). - VAE bakes into output quality -- the official SDXL VAE (madebyollin fp16 fix) is required to prevent washed-out colors. **Flux.1 [dev] / Flux.1 [schnell]:** - Transformer-based (not UNet). CLIP + T5-XXL dual text encoder architecture. - T5-XXL processes natural language with high fidelity -- full sentences work. Comma-separated tags are less optimal than in SD. - No attention weighting syntax in baseline Flux. Weights like `(term:1.5)` are either ignored or parsed inconsistently. - CFG guidance in schnell: 3.5-4.5. In dev: 3.5-7. Higher CFG than these ranges causes severe artifacts. - Very long prompts are handled better than SD, but subject/action/style ordering still matters. - Strong anatomy and hand quality at default settings compared to SD 1.x. **Midjourney v5/v6/v6.1:** - No explicit token limit, but prompt weighting uses `::` multi-prompt syntax and `--s` (stylize) control. - Earlier terms in the prompt carry more weight than later terms. - `--style raw` reduces aesthetic stylization. `--s 0` eliminates it almost entirely (often too flat). `--s 100-250` is the photorealism-to-artistic spectrum. - Parameter flags (`--ar`, `--v`, `--no`, `--cref`, `--sref`) are syntactically separate from the text prompt and must appear at the end. - v6+ responds to natural language far better than v5. v5-style comma-keyword prompts in v6 produce noticeably different (often worse) results. - `--no` is the primary negative mechanism. It is less precise than SD negative prompts but works for broad style exclusion. **DALL-E 3:** - Responds to natural language descriptions only. Tags and keyword syntax produce worse results. - ChatGPT/system prompt context affects output even when invisible to the user. - Has built-in safety and copyright filters that silently modify prompts ("prompt revision"). The user may not be prompting what they think they are prompting. - Does not support CFG, sampler, steps, seeds, or negative prompts. - If results seem randomly inconsistent, the system may be revising the prompt silently. Ask the user to check the "revised prompt" in the API response or UI tooltip. --- ### Step 3 -- Run the Seven-Category Diagnostic Evaluate the prompt against all seven categories systematically. Do not skip categories because they seem unlikely. **Category 1 -- Conflicting Descriptors** Look for pairs or groups of terms that encode opposing visual properties: - Medium conflicts: "watercolor" + "photorealistic" + "sharp focus" -- a medium (watercolor) that implies softness, combined with photographic precision - Lighting conflicts: "dramatic shadows" + "bright, even lighting" -- these are mutually exclusive lighting setups - Temporal conflicts: "sunset" + "golden hour" + "midday sun" -- pick one - Tone conflicts: "dark, gritty, noir" + "vibrant, colorful, cheerful" - Resolution conflicts: "painterly, impressionistic" + "4K, ultra-detailed, sharp" - Composition conflicts: "minimalist" + "highly detailed, intricate, busy" Each conflict forces the model to average two opposing instructions, producing a muddy or inconsistent result. Fix: identify the dominant intent, remove the contradicting term, and if both are genuinely desired (e.g., a detailed minimalist composition is actually possible), rewrite with precise spatial control. **Category 2 -- Style Paradigm Mismatch** The style descriptors do not belong to the same visual language family: - Mixing artistic-medium terms ("oil painting," "impasto texture") with photographic terms ("85mm lens," "bokeh," "ISO 800") -- these belong to incompatible visual paradigms - Mixing era/movement terms without coherence: "Art Deco Baroque cyberpunk" -- three unrelated style languages - Using film/photography jargon on a model that was not trained on photographic data (some SD 1.5 anime checkpoints do not respond to camera specs at all) - Using anime quality tags ("best quality, masterpiece") on photorealistic checkpoints -- these terms activate anime training data, not photo training data Fix: map all style descriptors to a single visual paradigm. Choose medium OR photography OR illustration OR 3D render, then use only the vocabulary of that paradigm. **Category 3 -- Underspecified Composition** The prompt lacks the spatial and structural cues needed to constrain the layout: - Subject with no positioning: "a warrior" -- foreground or background? Facing camera or turned? Close-up or full-body? - No depth cues: no near/far relationship, no layering, no atmospheric perspective reference - No lighting source: models default to flat ambient lighting when no light source is specified - No frame: "a forest" could be an aerial shot, a ground-level path view, a canopy shot, or a clearing -- the model picks randomly - Ambiguous subject count: "wolves in the snow" -- one wolf? Three? A pack of twenty? Severity: medium-to-high for subjects. Low for abstract or texture-only generations. Fix: add the minimum compositional triangle: (1) viewpoint/angle, (2) subject scale/framing, (3) light source. **Category 4 -- Overcrowded Token Budget / Adjective Flooding** Too many descriptors that compete for influence: - More than 8-12 meaningful descriptors in an SD 1.5 prompt exceeds the useful density threshold - Subjective adjectives ("beautiful," "stunning," "amazing," "perfect," "incredible") consume tokens without adding visual information -- they are evaluative, not descriptive - Redundant synonyms: "majestic, epic, grand, imposing" all instruct the model toward the same quality -- one strong term is more effective - In Midjourney, long adjective chains reduce the coherence of the core subject Token budget rule of thumb for SD: every comma-separated term uses approximately 1-3 tokens. A 30-word prompt may already be 50-60 tokens, leaving little budget for quality tags, style, and negative guidance. Fix: apply the "cut-by-half" test -- remove every term that does not directly describe a visible, physical quality. If the image would look the same without it, cut it. **Category 5 -- Negative Prompt Issues (SD-specific)** Negative prompts introduce their own failure modes: - Over-long negatives (100+ tokens) cut into the effective influence of the positive prompt because total attention is shared - Generic mega-negatives copied from the internet ("worst quality, low quality, normal quality, lowres, bad anatomy, bad hands, missing fingers, extra digit, fewer digits, cropped, jpeg artifacts, signature, watermark, username, artist name, blurry, text, error") applied to every prompt regardless of subject -- for a landscape, "bad anatomy" and "missing fingers" are irrelevant and waste token budget - Negating the subject accidentally: a portrait prompt with "face" in the negative will suppress facial features - Negating style you want: "painting" in the negative when the prompt asks for an "oil painting" style - Contradictory negatives: "no people" in the negative when the positive prompt features a character Fix: audit every term in the negative prompt against the positive. Remove any term that could match a desired element. Keep negatives focused and relevant to the specific content type. **Category 6 -- Parameter Misconfiguration** Settings interact with the prompt in ways that can override even well-written text: - **CFG scale:** - SD 1.5: optimal 6-9. Below 4: prompt ignored, random output. Above 14: oversaturated colors, burned highlights, artifacts. - SDXL: optimal 6-8. More sensitive than SD 1.5 -- 10+ causes noticeable degradation. - Flux.1 dev: 3.5-7. Flux.1 schnell: 1-3 (it is a distilled model; standard CFG ranges destroy it). - Midjourney: no direct CFG control. Use `--s` (stylize) instead. - **Step count:** - SD 1.5 with DPM++ 2M Karras: 20-30 steps is sufficient. Above 40 yields diminishing returns. Below 15 produces unfinished/blurry output. - SDXL: 25-35 steps optimal. Below 20 causes composition collapse. - LCM/Turbo distilled models: 4-8 steps. Running 30 steps on LCM oversmooths and degrades quality. - Flux.1 schnell: 4 steps is by design. Running more steps does not improve quality. - **Sampler:** - DPM++ 2M Karras: most consistent for photorealistic content, SD 1.5 and SDXL. - Euler a: more creative/varied but less consistent across seeds. - DDIM: good for inpainting, lower quality for full generations. - UniPC: fast with reasonable quality, good for testing iterations. - LCM: fast distilled sampler -- only for LCM LoRA or LCM checkpoint models. - **Resolution mismatch:** - SD 1.5 trained at 512x512. Generating at 768x768 directly causes anatomical issues (extra limbs, head duplication). Use hires.fix or img2img upscaling at 0.4-0.5 denoising strength. - SDXL trained at 1024x1024. Generating at 512x512 causes composition and anatomy problems. Never go below 768px on either dimension. - Generating non-square with incorrect orientation encoding: SDXL and SD 3 use target_size and crops_coords_top_left conditioning -- generating 1024x1024 but setting target_size to 768x768 produces incorrect scale rendering. **Category 7 -- Weight and Emphasis Imbalance (SD-specific)** Prompt weighting syntax must be used conservatively: - `(term:1.0)` -- no change. `(term:1.2)` -- modest emphasis. `(term:1.5)` -- strong emphasis, approaching artifact risk. `(term:2.0)` -- very likely to cause artifacts, bleed, or hallucinations. - Stacking weights additively: `((term))` in some implementations equals `(term:1.21)`. `(((term)))` equals approximately `(term:1.33)`. More than two stacked parentheses approaches artifact territory. - Competing high-weight terms: `(red hair:1.5), (blue eyes:1.5), (green dress:1.5)` forces the model to simultaneously oversaturate three competing attributes -- color artifacts result. - Under-weighting the subject while over-weighting modifiers: `a woman, (detailed background:1.5), (intricate texture:1.5)` can cause the background to dominate and the subject to become unclear. - SDXL: the dual encoder means weights propagate differently. Keep weights between 0.8 and 1.3 for reliable results. Fix: use weighting sparingly, only for the one or two most important elements. Reduce conflicting high weights. Use negative prompt suppression rather than weight reduction below 0.5. --- ### Step 4 -- Rank Issues by Severity After completing all seven category checks, rank every identified issue: - **High severity** -- This issue alone could fully explain the gap between expected and actual output. Fix this first. Example: "magical" triggering fantasy rendering when the user wanted photorealism. - **Medium severity** -- This issue contributes to the problem but may not be the sole cause. Fix after the high-severity issue is resolved. Example: missing camera specs that would reinforce a photorealistic intent. - **Low severity** -- This issue is a quality improvement rather than a fix. Address after the primary problems are resolved. Example: redundant synonyms consuming token budget without causing visible failure. State explicitly when you cannot determine severity without seeing the actual generated image. Ask the user for an image description if needed. --- ### Step 5 -- Write the Revised Prompt Apply all high and medium severity fixes. Apply low-severity fixes only if they do not change the user's intent. - Preserve the user's core subject and intent absolutely. If they wanted a mountain landscape, the revised prompt must still be a mountain landscape. - Make all changes minimally -- do not rewrite the prompt from scratch when targeted surgery is possible. - Front-load the most critical terms: subject first, then style, then composition, then quality modifiers. - For SD: quality tags first, then subject, then style, then composition, then technical specs (camera/lighting), then additional detail tags. - For Midjourney v6+: subject description first (in natural language), then style qualifiers, then technical parameters as flags at the end. - For Flux.1: full natural language sentence describing subject, action, environment, lighting, then style and technical references. - For DALL-E 3: complete natural language paragraph. No tags. Active voice ("a photographer captures...") outperforms passive ("a photograph of..."). - Document every single change separately. Do not bundle multiple changes into one table row. --- ### Step 6 -- Verify Model Limitations Are Not the Root Cause Some problems cannot be fixed with prompt changes. These must be identified honestly rather than routing the user toward an impossible prompt fix: - **Exact legible text in images** -- All diffusion models hallucinate text. DALL-E 3 handles short strings (1-3 words) in some contexts, but complex text is unreliable across all models. Honest fix: generate the image without text, add text in post using Photoshop, Canva, or similar. - **Precise spatial relationships** -- "Put the red ball exactly 3 centimeters to the left of the blue cube" is not achievable through prompting alone. Use ControlNet (for SD) or reference images (for MJ `--cref`/`--sref`) for precise layout. - **Consistent faces across multiple generations** -- Diffusion models do not have memory. Consistent character identity requires ControlNet face reference, IP-Adapter, Midjourney `--cref`, or fine-tuning (LoRA/DreamBooth). This is not fixable with prompting. - **Multiple distinct characters in a scene** -- Models frequently merge attributes of two characters (one character gets both their hair colors, their clothing blends). Mitigations exist (ControlNet multi-person, regional prompting in SDXL) but are workflow changes, not prompt changes. - **Hands and fingers** -- A structural limitation of most diffusion models due to training data distribution and the high variability of hand poses. Mitigation: add "perfect hands, correct anatomy, eight fingers, two thumbs" to positive; "deformed hands, extra fingers, missing fingers, malformed hands, mutated hands" to negative; use 30+ steps; CFG 6-8; consider ControlNet handpose reference. Set honest expectations: this reduces but does not eliminate hand errors. - **Exact color matching to a hex code** -- Models do not process hex values. Describe colors in natural language with multiple reference points: "deep burgundy red, the color of aged Merlot wine" works better than "#7B2240." --- ### Step 7 -- Compile the Full Debug Report Assemble all findings into the structured output format. Include: - Original prompt, exact - All issues detected, categorized and severity-rated - A diagnostic explanation in natural language that explains the causal chain - The revised prompt, exact - A changes table with every modification documented - Numbered follow-up recommendations if the revised prompt still fails - Prevention tips tailored to the specific mistake pattern found --- ## Output Format ``` ## Prompt Debug Report ### Original Prompt [user's exact prompt, reproduced verbatim -- do not paraphrase] Model: [model name and version] | Parameters: [CFG: X | Steps: X | Sampler: X | Resolution: XxX | Seed: X or "not provided"] Negative Prompt (if provided): [exact negative prompt text, or "none provided"] --- ### Issues Detected | # | Category | Severity | Specific Problem Found | |---|-------------------------------|----------|-----------------------------------------------------------------| | 1 | [Category name] | HIGH | [Specific conflicting/problematic terms cited verbatim] | | 2 | [Category name] | MEDIUM | [Specific issue with quoted term from the prompt] | | 3 | [Category name] | LOW | [Specific minor issue with quoted term] | Severity key: HIGH = primary cause | MEDIUM = contributing factor | LOW = quality improvement --- ### Diagnosis **Primary failure mode:** [1-sentence summary of the main cause] [Paragraph 1: Explain the causal chain -- which specific terms are causing which specific visual outcomes, and why. Reference the model's architecture where relevant. Be specific about term interactions.] [Paragraph 2: Secondary contributing factors. Note any parameter settings that compound the prompt issues. Note any model limitations that apply.] --- ### Revised Prompt [corrected prompt, exact and complete, ready to copy-paste] Model: [model] | Parameters: [CFG: X | Steps: X | Sampler: X | Resolution: XxX] Revised Negative Prompt (SD only, if applicable): [revised negative prompt] --- ### Changes Made | # | Original Term or Setting | Revised To | Reason |
Voir sur GitHub
Ce SKILL.md est tres volumineux, SkillsMP affiche donc ici seulement la premiere section. Voir sur GitHub