Skip to main content

ai-image-style-transfer

Identifies a visual style's defining characteristics and translates them into AI image generation trigger words for Midjourney, DALL-E, and Stable Diffusion with model-specific syntax. Use when the user asks to replicate a visual style in AI art, transfer a style between images, describe an art style for AI generation, or match a reference image's aesthetic. Do NOT use for basic prompting in a single model (use the model-specific prompting skill), maintaining character consistency (use midjourney-consistency), or translating complete prompts between models (use prompt-translation).

Informations de source

Dépôt
FerroxLabs/murage
Dernière activité de la source
1 septembre 2026 à 13:26
Langue détectée de SKILL.md
anglais
Étoiles
9
Forks
2

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Explorateur de fichiers
2 fichiers

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
ai-image-style-transfer
description
Identifies a visual style's defining characteristics and translates them into AI image generation trigger words for Midjourney, DALL-E, and Stable Diffusion with model-specific syntax. Use when the user asks to replicate a visual style in AI art, transfer a style between images, describe an art style for AI generation, or match a reference image's aesthetic. Do NOT use for basic prompting in a single model (use the model-specific prompting skill), maintaining character consistency (use midjourney-consistency), or translating complete prompts between models (use prompt-translation).
license
Apache-2.0
metadata
{"author":"foundry-skills","version":"1.0.0","tags":"ai-image-generation design analysis","category":"design-creative","subcategory":"ai-image-generation","depends":"","disclaimer":"none","difficulty":"intermediate"}
# AI Image Style Transfer ## When to Use **Use this skill when:** - The user wants to replicate a named visual style (Impressionism, Art Nouveau, Bauhaus, cyberpunk, vaporwave) across one or more AI generation models and needs model-specific prompt syntax - The user has a reference image and wants to capture its aesthetic in a new generation -- they need the style decomposed, not just described back to them - The user's style prompts are producing generic or inconsistent results and they need targeted diagnosis (wrong trigger words, style fighting the subject, stylize parameter too low) - The user wants to understand which visual properties of a complex style are achievable in AI generation vs. which require post-processing - The user wants to combine two or more distinct styles (e.g., "gothic Art Nouveau," "Soviet constructivist cyberpunk") and needs a principled blending strategy - The user wants to target a specific medium-style (film photography aesthetics, woodblock print, risograph printing, linocut) and needs both the visual decomposition and the specific model vocabulary - The user is building a consistent visual identity across multiple AI generations and needs repeatable style tokens, not one-off prompts **Do NOT use when:** - The user wants a complete prompt for a single model without needing style transfer analysis -- use `midjourney-prompting`, `dalle-prompting`, or `stable-diffusion-prompting` - The user wants to keep a specific character's face, outfit, or identity consistent across multiple scenes -- use `midjourney-consistency` - The user has a complete working prompt in one model and wants it converted to another model's syntax verbatim -- use `prompt-translation` - The user is asking about AI model capabilities in general, not prompting -- use a general AI explainer skill - The user wants photography retouching or style application in Lightroom/Photoshop, not AI generation -- use a photo editing skill - The user is asking about copyright or the ethics of style replication -- this skill focuses on technical execution, not policy --- ## Process ### Step 1: Gather Style Inputs Before decomposing anything, collect the minimum required information. Missing any of these leads to incorrect prompt architecture. - **Style source** -- Get the most precise form available. A named art movement (Post-Impressionism) is less precise than a named technique (Pointillism). A reference image is more precise than either. An era description ("1970s sci-fi paperback cover art") is a valid starting point but requires inference. - **Target model(s)** -- Ask explicitly: Midjourney (and which version: v6.1, v7, Niji 6?), DALL-E 3, or Stable Diffusion/ComfyUI/Flux? Each has fundamentally different prompt architectures. If the user wants all three, note that each requires a separately optimized prompt -- not a copy-paste. - **Subject matter** -- What is the content being rendered in this style? Style interacts with subject: a portrait in Fauvism behaves differently than a landscape in Fauvism because color distortion reads differently on faces. - **Fidelity level** -- Is the user aiming for "unmistakably this style" (high fidelity, style dominates) or "loosely inspired by" (style as flavor, subject leads)? This determines stylize parameter values and how aggressively to weight style tokens. - **Output use case** -- Print, digital display, social media post, concept art reference? This affects aspect ratio, resolution guidance, and post-processing recommendations. ### Step 2: Decompose the Style into the Six Visual Properties Every visual style is a combination of these six properties. Work through all six systematically. Skipping one causes incomplete prompts that produce generic-looking results. **Property 1: Color Palette** - Identify the dominant color temperature: warm (sunset orange, ochre, gold), cool (slate blue, steel grey, ice white), or neutral (desaturated, earth tones, sepia) - Name specific hues using art vocabulary: cerulean, vermilion, viridian, raw umber, cadmium yellow, Prussian blue -- these terms appear in training data and carry more weight than "red" or "blue" - Identify contrast behavior: high contrast (dark darks, light lights with no midtones) vs. tonal range (smooth gradation across all values) vs. low key (predominantly dark) vs. high key (predominantly light) - Note any signature color behavior: duotone, triadic color schemes, split complementary, analogous harmony, or intentional color dissonance (Expressionism) - Flag if the palette has a printing or medium origin: risograph uses two or three overlapping spot colors with halftone dot patterns; woodblock prints show ink bleed at edges; Polaroid shifts toward warm green-yellow shadows **Property 2: Line Quality and Edge Behavior** - This is the property most beginners omit. Every style has a characteristic relationship between shapes: hard graphic edges (Art Deco, Pop Art), soft blended edges (Impressionism, Sfumato), visible drawn outlines (cel animation, comic books), no visible outlines (plein air painting, photography) - Characterize line weight if present: uniform weight (technical illustration), variable weight calligraphy (traditional ink painting), tapered brush lines (manga), thick contour with no detail interior (bold graphic style) - Identify mark-making vocabulary if relevant: stipple (Pointillism, engraving), cross-hatching (pen-and-ink illustration), impasto drag marks (palette knife painting), dry brush texture (traditional Chinese ink wash), gestural swirls (Van Gogh's Post-Impressionist language) **Property 3: Texture and Surface Quality** - Identify the implied physical substrate: canvas grain, paper tooth, glass smoothness, concrete roughness, film grain, digital smoothness - Note texture behavior: does it appear uniformly across the image or only in specific areas (e.g., visible brushwork in shadows but smooth in highlights)? - Quantify grain if applicable: fine film grain (ISO 400 equivalent) vs. heavy grain (ISO 3200 push-processed) vs. halftone dots (offset printing screen at 85 LPI) vs. noise (digital sensor noise) - For printed/mechanical styles: risograph mottling, letterpress impression depth, screen printing misregistration, linocut uneven ink coverage -- these have specific vocabulary that AI models recognize **Property 4: Composition and Spatial Logic** - Identify the depth model: true 3D perspective with depth recession, shallow depth (Japonisme, Art Nouveau), flat 2D with no depth cues (Matisse cutouts, Swiss design), isometric (pixel art, technical illustration), forced perspective (propaganda posters) - Note focal point convention: centered and symmetrical (Art Deco, Byzantine iconography), rule-of-thirds dynamic (photojournalism, landscape painting), all-over composition with no single focal point (Abstract Expressionism, pattern design) - Identify compositional framing devices used by this style: decorative borders (Art Nouveau vine frames, medieval illumination), bleed-to-edge (fashion photography), negative space as a compositional element (minimalism, Japanese sumi-e), overlapping planes (Cubism, collage) - Note figure-ground relationship: high figure-ground separation (poster art), ambiguous figure-ground (Escher, Op Art), ground as important as figure (landscape tradition) **Property 5: Lighting Model** - Identify the light source type: single strong directional (Caravaggio, film noir), diffuse ambient (overcast plein air), multiple sources (studio photography, three-point lighting), no implied light source (flat illustration, stained glass) - Characterize shadow behavior: cast shadows vs. form shadows vs. both; hard-edged shadows (high contrast, noon light, harsh artificial) vs. soft gradated penumbra (window light, cloudy sky) - Note signature lighting effects: chiaroscuro (90% shadow, dramatic value jump), rim lighting (backlit silhouette effect), bioluminescence, neon glow with light bloom, golden hour color shift (everything shifts +15° hue toward orange), underwater caustic patterns - For photography styles: specify the quality of light: Rembrandt lighting (triangle highlight on shadow side cheek), split lighting (50/50 face divide), butterfly lighting (shadow under nose), catchlights in eyes **Property 6: Rendering and Realism Level** - Place the style on the realism spectrum: photorealistic (no visible abstraction), hyperreal (more detailed than perception), stylized realistic (recognizable but simplified), semi-abstract (forms suggested), fully abstract (no representational content) - Identify the detail distribution: hyperdetailed everywhere (maximalism, Baroque), detailed foreground with loose background (impressionist sketch), flat fill with precise edge detail only (stained glass, cloisonné), overall simplification (icon design, Shaker aesthetic) - Note perspective handling: natural single-point perspective, two-point perspective (architectural rendering), three-point (dramatic up/down angles), fish-eye (180° distortion), orthographic (no perspective convergence) - Identify stylistic distortions if they are characteristic: elongation (El Greco, Modigliani, Art Nouveau figures), geometric fragmentation (Cubism), proportion exaggeration (chibi, caricature), impossible geometry (Surrealism) ### Step 3: Identify Style-Defining Differentiators Most styles share surface similarities with adjacent styles. Identify 2-4 properties that specifically distinguish this style from the most commonly confused alternatives. This step prevents the most common failure mode: generating a style that looks like something adjacent. - Impressionism vs. Post-Impressionism: Impressionism uses loose, small dabs of unmixed color to capture light and atmosphere; Post-Impressionism retains that texture but reintroduces structured composition and symbolic color (Cézanne's geometric planes, Van Gogh's swirling directional marks) - Art Nouveau vs. Art Deco: Art Nouveau uses organic flowing curves derived from plant and natural forms, asymmetric compositions, and muted earth/jewel tones; Art Deco uses geometric rectilinear forms, strong symmetry, and high-contrast metallic palettes - Vaporwave vs. Synthwave vs. Lo-fi: Vaporwave uses classical marble statues, Roman columns, early CGI renders, oversaturated pink/purple/cyan, glitch artifacts, and soft VHS noise; Synthwave uses retro-futurist 80s neon grids, chrome lettering, lens flare, and dramatic sunset gradients; Lo-fi uses muted warm tones, analog grain, and cozy domestic scenes without the digital artifact vocabulary - Anime styles: distinguish between shojo (delicate features, large eyes, soft pastel watercolor backgrounds, flower motifs), shonen (bold lines, dynamic motion blur, high contrast, muscular forms), and mecha (hard-surface rendering, technical detail, dramatic perspective, chrome and shadow) ### Step 4: Map Properties to Model-Specific Vocabulary Each model has different mechanisms for style activation. Using the wrong vocabulary for a model produces weak or ignored style signals. **Midjourney (v6.1 and v7) Vocabulary and Parameters:** - Style tokens work best as specific noun phrases placed early in the prompt: "oil painting," "woodblock print," "gouache illustration," "linocut" -- not adjectives like "painterly" or "illustrative" - The `--stylize` parameter (range 0-1000, default 100) is the most powerful style fidelity control. Values of 250-400 produce strong aesthetic interpretation. Values above 600 let Midjourney's aesthetic model dominate the reference style. Values below 50 produce literal but aesthetically flat results. For style transfer, 200-500 is the working range. - `--style raw` disables Midjourney's default beautification engine. Use this for styles that should look rough, aged, or low-tech (Polaroid, zine aesthetics, risograph, vintage poster) - `--sref [URL]` (style reference) in v6.1 and v7 extracts color palette and surface texture from a reference image at weight 0-1000. At weight 100 (default), it's subtle. At 500, it dominates. Use `--sw 500` to boost style reference weight. Note: `--sref` captures palette and texture but not composition -- it does not impose the reference image's layout - `--niji 6` is a separate model tuned for anime/illustration styles. For any cel-animated, manga, or anime-adjacent style, switch to `--niji 6` instead of `--v 6.1` - Aspect ratio `--ar` affects how composition directs: `--ar 2:3` activates portrait compositional conventions, `--ar 16:9` activates cinematic/landscape conventions, `--ar 1:1` activates centered symmetrical tendencies - The `--no` parameter is a hard exclusion: `--no photograph, realistic, 3D render` are standard additions for painterly/illustrative styles **DALL-E 3 Vocabulary:** - DALL-E 3 processes natural language sentences, not keyword tags. Style must be described in prose: "painted with thick impasto strokes of unmixed color, with a rough canvas texture visible throughout" outperforms "impasto painting" - Medium specification is critical: explicitly name "oil paint," "watercolor on cold-press paper," "digital illustration," "gouache on toned paper" -- DALL-E 3 has strong training on medium vocabulary - DALL-E 3 responds well to art historical references when framed as technique descriptions: "in the manner of 19th-century French academic painting, with smooth blended skin tones, glazed luminosity, and dramatic chiaroscuro" is effective; naming specific living artists is not permitted and naming deceased artists may produce inconsistent results -- technique descriptions are more reliable - Negative prompting is not directly supported in DALL-E 3's API/interface. Instead, use affirmative language to crowd out unwanted qualities: "flat graphic style with no photographic realism" excludes photorealism without a negative keyword - DALL-E 3 handles complex compositional instructions better than other models. Describe spatial arrangement explicitly: "the figure occupies the left third of the frame, with a receding landscape in the right two-thirds using atmospheric perspective" **Stable Diffusion / Flux / ComfyUI Vocabulary:** - SD prompt architecture uses weighted tags: `(tag:weight)` where 1.0 is default, 1.3 is strong emphasis, 1.5 is very strong, and above 1.5 begins to produce distortion. Common style weights: `(oil painting:1.3)`, `(art nouveau:1.2)`, `(impressionist:1.2)` - Leading quality tokens still matter on base models: `masterpiece, best quality, highly detailed` placed first. On SDXL and Flux, these matter less but still provide a quality floor - Negative prompts are critical for style isolation. Standard style-contamination exclusions: `photorealistic, 3d render, cgi, digital art` (when targeting traditional media), or `painting, brushstrokes, traditional media` (when targeting photography) - LoRA (Low-Rank Adaptation) files provide the most accurate style replication in SD. Syntax: `<lora:filename:weight>` where weight 0.6-0.8 is the typical working range. Above 0.9 often causes saturation and artifact buildup. Multiple LoRAs can be stacked but total weight should not exceed 1.5 combined. Always note that LoRAs must be downloaded and installed locally -- they cannot be assumed to be available - Textual Inversion embeddings (`.pt` or `.bin` files) also carry style vocabulary: `embedding_name` placed in the prompt activates the concept. Embeddings are generally subtler than LoRAs. - Flux.1 models (dev, schnell) use natural language closer to DALL-E 3 than traditional SD tag syntax. With Flux, write descriptive sentences while retaining some key style tags: "impressionist oil painting with visible short brushstrokes and warm dappled light, (Monet-style:1.1)" - CFG scale (classifier-free guidance) affects style fidelity: 7-9 is standard for most styles. For very specific style lock-in, 9-11 strengthens prompt adherence but reduces variation. Below 5 produces dreamlike but style-loose results. ### Step 5: Diagnose Style Conflicts Identify any terms in the proposed subject description that will fight the target style. Style conflict is the most common reason prompts fail. - Photographic subject terms conflict with painterly styles: "a woman" is neutral, but "a woman standing in a room" begins to pull toward photographic realism. Add "painted portrait of a woman" to anchor the medium - Specific named locations (Eiffel Tower, Times Square) pull toward photographic documentary in all models. Pair them with the style medium to counteract: "oil painting of the Eiffel Tower in the Impressionist style, not photographic" - Highly detailed technical subjects (circuit boards, mechanical gears) pull SD and MJ toward technical illustration regardless of style. Weight the style term higher when subjects are technical: `(impressionist painting:1.5), circuit board` - Fantasy/sci-fi subjects (dragons, spaceships) pull MJ toward its default "fantasy art" aesthetic. Use `--style raw` to reduce this and let the specified style dominate - Color-heavy subjects can override a muted palette. If the target style uses muted earth tones but the subject is "a rainbow," acknowledge the conflict and suggest either adapting the subject description or accepting palette compromise ### Step 6: Assess Transfer Fidelity by Property Be honest and specific about what each model can and cannot do with each style property. Assign each property a fidelity rating: **High**, **Medium**, or **Low**, with a specific reason. - **High**: The property transfers reliably and consistently. The model has strong training signal for this exact visual attribute. Example: Midjourney reliably transfers color palettes because it has strong color following. - **Medium**: The property transfers partially or inconsistently. Results require iteration (try the prompt 4-8 times and select). Example: Composition symmetry is approximate in all models -- slight asymmetries are common. - **Low**: The property is suggested at best. The AI frequently drifts from the target. Acknowledge this explicitly and recommend alternatives (LoRA, post-processing, compositional templates). Example: Exact brushwork direction (Van Gogh's swirling marks) is Low fidelity in DALL-E 3. ### Step 7: Compile the Complete Style Transfer Specification Assemble all outputs into the standard format below. Include: - The full style analysis table - Model-specific prompts with all required syntax - Key trigger terms called out explicitly - Transfer fidelity table - Honest limitations with actionable workarounds --- ## Output Format ```markdown ## Style Transfer: [Style Name] ### Style Differentiators *What distinguishes this style from commonly confused alternatives:* - [Differentiator 1: "Unlike X, this style uses Y because Z"] - [Differentiator 2] - [Differentiator 3] ### Style Analysis | Property | Visual Description | Key Descriptors for Prompts | |--------------------|------------------------------------------------------------|----------------------------------------------------------| | Color Palette | [Specific hues, temperature, contrast behavior] | [Art vocabulary terms for prompts] | | Line & Edge | [Edge quality, line weight, outline behavior] | [Specific line/edge vocabulary] | | Texture & Surface | [Physical substrate, mark-making, grain/print artifacts] | [Texture vocabulary] | | Composition | [Depth model, focal point, spatial logic] | [Compositional terms] | | Lighting | [Source type, shadow behavior, signature effects] | [Lighting vocabulary] | | Rendering Level | [Realism spectrum, detail distribution, distortions] | [Rendering vocabulary] | ### Style Conflict Warnings *These subject or prompt terms will fight this style -- avoid them:* - [Conflict term 1]: [Why it conflicts and what to use instead] - [Conflict term 2]: [Why it conflicts and what to use instead] --- ### Model-Specific Prompts #### Midjourney (v6.1) ``` /imagine prompt: [subject], [medium/style anchor], [color palette description], [texture/line quality], [composition description], [lighting description], [rendering level], [2-3 reinforcing style adjectives] --ar [ratio] --v 6.1 --s [200-500] --style raw (if applicable) --no [conflict terms] ``` **Key trigger terms:** [list 4-6 terms that activate this style in MJ, with notes on why they work] **Parameter rationale:** `--s [value]` because [specific reason]; `--style raw` [use/omit] because [specific reason] **Style reference note:** [If applicable: how to use --sref for this style] --- #### DALL-E 3 ``` [2-4 sentence natural language prompt. First sentence: medium and style name. Second sentence: color palette and texture description. Third sentence: composition and lighting. Fourth sentence: what this image is NOT (photorealistic, etc.) using affirmative exclusion language.] ``` **Key trigger phrases:** [list 4-5 descriptive phrases that DALL-E 3 responds to for this style] **DALL-E 3 notes:** [Any specific handling notes: what this model does well vs. poorly for this style] --- #### Stable Diffusion (SDXL / Flux.1) ``` Positive prompt: masterpiece, best quality, [medium tag], [(style name:1.2)], [(color palette terms:1.1)], [texture/surface terms], [composition terms], [lighting terms], [(rendering terms:1.1)], [reinforcing style vocabulary], <lora:suggested_lora_name:0.7> Negative prompt: [terms that contaminate this style], lowres, bad anatomy, worst quality, blurry, [medium-specific exclusions] ``` **Key trigger tags:** [list with weights -- which tags carry the most style activation signal] **LoRA recommendation:** [Specific LoRA category to search for + typical weight range 0.6-0.8; note user must install locally] **CFG recommendation:** [value range + rationale] **Flux.1 variation:** [Modified syntax for Flux.1 if different from SDXL] --- ### Transfer Fidelity Assessment | Property | Midjourney | DALL-E 3 | SD/Flux | Specific Notes |
Voir sur GitHub
Ce SKILL.md est tres volumineux, SkillsMP affiche donc ici seulement la premiere section. Voir sur GitHub