| name | qwen-image-edit |
| description | Build Qwen Image Edit workflows covering model loading, conditioning, LoRAs, prompt patterns, and XY plot testing |
| globs | ["**/*.json"] |
Qwen Image Edit Workflows
Overview
Qwen Image Edit uses a vision-language model (Qwen2.5-VL) to edit images based on natural language instructions. The model "sees" the source image through CLIP conditioning and generates an edited version.
Models
Required Components
| Component | Node | Model Name | Notes |
|---|
| UNET | UNETLoader | qwen_image_edit_2511_bf16.safetensors | Official 2511 edit model (bf16) |
| CLIP | CLIPLoader (type=qwen_image) | qwen_2.5_vl_7b_fp8_scaled.safetensors | Shared across all Qwen models |
| VAE | VAELoader | qwen_image_vae.safetensors | Qwen-specific VAE |
Alternative UNET Models
| Model | Path | Focus |
|---|
qwenImageEditRemix_v10 | qwenImageEditRemix_v10.safetensors | Community remix, general editing |
qwenUltimateRealism_v11 | Qwen/imageized/qwenUltimateRealism_v11.safetensors | Product photography, hyper-realistic |
copaxTimeless | Qwen/realistic/copaxTimeless_qwenUltraRealistic.safetensors | Ultra-realistic portraits |
qwnImageEdit_v16Bf16 | Qwen/abliterated/qwnImageEdit_v16Bf16.safetensors | Abliterated (uncensored) |
Conditioning Nodes
TextEncodeQwenImageEditPlusAdvance_lrzjason (Recommended)
From the qweneditutils custom node pack. The Advanced variant is preferred because it:
- Outputs a LATENT directly (no need for separate EmptyLatentImage)
- Has separate VL-resize and non-resize image slots for fine control
- Supports target_size control for output resolution
- Includes a pad/center/disabled crop method with pad_info output
Required Inputs:
- clip: CLIP
- prompt: STRING — natural language edit instruction
Optional Inputs:
- vae: VAE — needed for image encoding and latent output
- vl_resize_image1-3: IMAGE — images that get VL-resized (downscaled for vision encoder)
- not_resize_image1-3: IMAGE — images kept at full resolution
- target_size: [1024, 1344, 1536, 2048, 768, 512] (default 1024)
- target_vl_size: [392, 384] (default 384)
- upscale_method: [lanczos, bicubic, area]
- crop_method: [pad, center, disabled]
- instruction: STRING — system instruction template (has sensible default)
Outputs (10):
[0] conditioning_with_full_ref: CONDITIONING — use as positive conditioning
[1] latent: LATENT — auto-scaled latent, feed directly to KSampler
[2] target_image1: IMAGE — processed target-size image
[3] target_image2: IMAGE
[4] target_image3: IMAGE
[5] vl_resized_image1: IMAGE — VL-resized version
[6] vl_resized_image2: IMAGE
[7] vl_resized_image3: IMAGE
[8] conditioning_with_first_ref: CONDITIONING — conditioning with only first ref
[9] pad_info: ANY — padding info for later unpadding
Key advantage: Output [1] (latent) eliminates the need for a separate EmptyLatentImage or VAEEncode node. The Advanced node handles latent creation internally at the correct resolution.
Other Conditioning Variants
- TextEncodeQwenImageEditPlus (Phr00t v2, built-in) is simpler: 4 image inputs, outputs only CONDITIONING. Requires separate EmptyLatentImage. Good for quick edits.
- TextEncodeQwenImageEditPlus_lrzjason: 5 image inputs, resize toggles, but less control than Advance
- TextEncodeQwenImageEditPlusPro_lrzjason: Per-image VL resize selection via
vl_resize_indexs string, main_image_index control
Lightning LoRAs (Fast Generation)
4-Step Lightning (2511 Edit)
{
"class_type": "LoraLoaderModelOnly",
"inputs": {
"model": ["<unet_node>", 0],
"lora_name": "Qwen-Image-Edit-2511-Lightning-4steps-V1.0-bf16.safetensors",
"strength_model": 1.0
}
}
Settings: steps=4, cfg=1.0, sampler=euler, scheduler=simple, denoise=1.0
4-Step Lightning (General Qwen)
For non-edit models (txt2img, 2512):
Qwen-Image-Lightning-4steps-V1.0.safetensors (strength 1.0)
8-Step Lightning
Qwen-Image-Lightning-8steps-V1.0.safetensors, higher detail than 4-step
Sampler Settings
| Preset | Steps | CFG | Sampler | Scheduler | Denoise | LoRA |
|---|
| Lightning 4-step (2511 edit) | 4 | 1.0 | euler | simple | 1.0 | 2511-Lightning-4steps |
| Lightning 8-step | 8 | 1.0 | euler | simple | 1.0 | Lightning-8steps |
| Standard edit | 40 | 4.0 | euler | simple | 0.75 | none |
| Quality edit | 50 | 4.0 | euler | simple | 0.5-0.8 | none |
The sub-1.0 denoise rows REQUIRE a VAEEncode latent. A denoise low enough to
shorten the sampling schedule — which 0.5-0.8 certainly is — keeps part of the
incoming latent, so that latent has to BE the source image. Wire latent_image from
a VAEEncode of the source (or from a node that emits a source-derived latent, like
TextEncodeQwenImageEditPlusAdvance_lrzjason output [1]). Pairing these rows with an
EmptyLatentImage runs clean and returns a flat, near-uniform field — an empty latent
has no source content to preserve. Feeding the reference through
TextEncodeQwenImageEditPlus does not rescue it: that image rides on CONDITIONING,
which steers denoising but never seeds the sampler's starting state.
Denoise for editing: Lower denoise = closer to source — provided the latent IS the source. 0.5-0.8 range for standard editing on a VAEEncode latent. Lightning uses 1.0 (model handles fidelity internally).
Resolutions
This table is for Qwen-Image TEXT-TO-IMAGE. Do not pick an edit graph's output
size from it. An edit graph's geometry is decided by the SOURCE image, not by you
— see "Resolution on an edit graph" below. Choosing 1104x1472 here for an edit was
#2681.
Qwen-Image operates at ~1.6 megapixels natively:
| Aspect | Resolution | Use Case |
|---|
| Square | 1328x1328 | General |
| Portrait 3:4 | 1104x1472 | Portraits |
| Portrait 9:16 | 928x1664 | Phone format |
| Landscape 4:3 | 1472x1104 | Landscape scenes |
| Landscape 16:9 | 1664x928 | Widescreen |
| Video-ready | 832x480 | For WAN 2.2 FLF pipeline |
For video pipelines: Use 832x480 to match WAN 2.2's default resolution.
Resolution on an edit graph
TextEncodeQwenImageEdit and TextEncodeQwenImageEditPlus do not take a size. They
scale every reference image to a hard-coded int(1024 * 1024) px — ~1.05 MP, at the
source's own aspect ratio — VAE-encode it, and hand it to the model as a reference
latent (comfy_extras/nodes_qwen.py).
The model then lays the reference tokens and the tokens it is generating on one
shared, centred coordinate grid (comfy/ldm/qwen_image/model.py, process_img), so
reference position (i, j) and output position (i, j) mean the same place only when the
two grids are the same size. That is the whole reason the official templates run the one
image through FluxKontextImageScale and then feed the sampler a VAEEncode of that
scaled image — both branches then see the same pixels at the same ~1 MP scale and the
same aspect. Every PREFERRED_KONTEXT_RESOLUTIONS entry is ~1.05 MP for the same reason.
It is agreement to within the encoder's round-to-8, not exact equality, and the
difference is worth knowing precisely. FluxKontextImageScale snaps to a preferred pair;
the encoder then renormalises that to its own 1,048,576 px budget. For 12 of the 18
preferred pairs the two land on the same latent grid. For the other 6 — 688x1504,
800x1328, 832x1248 and their landscape mirrors — the reference lands one latent row or
column off: at 800x1328 the sampler's grid is 100x166 and the reference's is 99x165.
ComfyUI's own bundled 2511 template does exactly this, so a sub-patch offset is evidently
fine in practice. The failure this page is about is one of SCALE, not of rounding — an
empty latent at 1104x1472 sits 1.24x away linearly, not one row.
So on an edit graph you do not choose a resolution — you inherit one:
- Right:
LoadImage -> FluxKontextImageScale -> (TextEncodeQwenImageEditPlus
and VAEEncode) -> KSampler latent_image.
- Wrong: an
EmptyLatentImage at a size from the table above. Its dimensions are
literals; the reference's are computed from the source when the graph runs. At
1104x1472 (1.63 MP) against a 1.05 MP reference the grids are 1.24x apart linearly,
the model cannot copy detail across them, and it re-synthesises the subject instead —
materials come back looking plastic/CGI and printed detail comes back as a generic
shape (#2681). Nothing errors; the image just is not the edit you asked for.
create_workflow (action:"validate") now flags this pairing as
edit_reference_empty_latent.
Prompt Patterns
Edit Instructions (Natural Language)
"Change the black cat into a cute girl with a black bodysuit and jeans"
"Make the sky a dramatic sunset with orange and purple clouds"
"Add a red sports car parked in front of the house"
"Remove the person on the left and fill with the background"
Multi-Angle LoRA (qwen-image-edit-2511-multiple-angles-lora)
Uses <sks> token with structured angle/distance prompts:
<sks> front view eye-level shot close-up
<sks> front-right quarter view low-angle shot medium shot
<sks> back view elevated shot wide shot
Template: <sks> {direction} view {angle} shot {distance}
Directions: front, front-right quarter, right side, back-right quarter, back, back-left quarter, left side, front-left quarter
Angles: low-angle, eye-level, elevated, high-angle
Distances: close-up, medium shot, wide shot
Negative Conditioning
Always use ConditioningZeroOut for negative conditioning with Qwen edit:
{
"class_type": "ConditioningZeroOut",
"inputs": { "conditioning": ["<positive_cond_node>", 0] }
}
Complete Workflow: Lightning Edit (Advanced Node)
Uses TextEncodeQwenImageEditPlusAdvance_lrzjason, which outputs the latent directly, so no EmptyLatentImage is needed.
{
"1": { "class_type": "UNETLoader", "inputs": { "unet_name": "qwen_image_edit_2511_bf16.safetensors", "weight_dtype": "default" }},
"2": { "class_type": "LoraLoaderModelOnly", "inputs": { "model": ["1", 0], "lora_name": "Qwen-Image-Edit-2511-Lightning-4steps-V1.0-bf16.safetensors", "strength_model": 1 }},
"3": { "class_type": "CLIPLoader", "inputs":
Key connections:
"latent_image": ["6", 1]: KSampler gets its latent directly from the Advanced node's output [1]
"positive": ["6", 0]: conditioning_with_full_ref from output [0]
"vl_resize_image1": ["5", 0]: source image goes into VL-resize slot (downscaled for vision encoder)
Simpler Alternative (Phr00t v2)
If qweneditutils custom node is unavailable, use the built-in TextEncodeQwenImageEditPlus with a separate EmptyLatentImage:
{
"6": { "class_type": "TextEncodeQwenImageEditPlus", "inputs": {
"clip": ["3", 0], "prompt": "<edit instruction>", "vae": ["4", 0], "image1": ["5", 0]
}},
"8": { "class_type": "EmptyLatentImage", "inputs": { "width": 1024, "height": 1024, "batch_size"
Replace node 6 and add node 8. KSampler latent_image connects to ["8", 0] instead of ["6", 1].
Do not leave that EmptyLatentImage wired to the sampler. It is shown above only
because it is what the Plus encoder's own signature leaves you needing, and it fails in a
different way on each side of denoise 1.0:
- Below 1.0 it renders a flat, near-uniform field, always. The encoder emits CONDITIONING
only, so the latent is genuinely empty, and a truncated sigma schedule exists to
preserve an incoming latent that here has nothing in it (#2678).
- At 1.0 it renders a plausible image that is not your source — unless its width and
height happen to equal the geometry the encoder derived, which is ~1.05 MP at the
source's aspect ratio. A 1024x1024 empty latent over an exactly-square source does line
up, and is fine. 1104x1472 over that same source does not: the model aligns reference
and output on one shared grid, cannot copy across grids that far apart, and
re-synthesises instead (#2681). See "Resolution on an edit graph".
The second case is the trap, because it depends on a source you may not have looked at
and it fails silently. A VAEEncode removes the coincidence — it cannot be the wrong
size, because it is derived from the same pixels the encoder saw.
Feed latent_image from a VAEEncode of the same image you gave the encoder, scaled
once up front so both branches see the same pixels at the same scale:
{
"5b": { "class_type": "FluxKontextImageScale", "inputs": { "image": ["5", 0] }},
"6": { "class_type": "TextEncodeQwenImageEditPlus", "inputs": {
"clip": ["3", 0], "prompt": "<edit instruction>", "vae": ["4", 0], "image1": ["5b", 0]
}}
KSampler latent_image connects to ["8", 0]. Any denoise is then meaningful: 1.0 for
a full edit, sub-1.0 to stay closer to the source.
Basic Variant (Official ComfyUI Example)
The official "Qwen 2511 Edit Simple" example uses newer built-in nodes for model patching and image scaling:
Additional nodes in the official pipeline:
ModelSamplingAuraFlow (shift=3.1): Flow matching shift applied to the UNET. Used instead of ModelSamplingSD3.
CFGNorm (strength=1): Normalizes CFG guidance for more stable generation. Applied after ModelSamplingAuraFlow.
FluxKontextImageScale: Auto-scales input images to the correct resolution for Qwen. No manual size parameters needed.
FluxKontextMultiReferenceLatentMethod (method=index_timestep_zero): Applied to both positive and negative conditioning. Handles multi-reference latent indexing.
VAEEncode: Encodes the scaled image to latent (instead of EmptyLatentImage).
Official pipeline flow:
UNETLoader → [LoraLoaderModelOnly] → ModelSamplingAuraFlow (shift=3.1) → CFGNorm (strength=1) → MODEL
CLIPLoader (qwen_image) → CLIP
VAELoader → VAE
LoadImage → FluxKontextImageScale → scaled_image
├─ TextEncodeQwenImageEditPlus (positive) → FluxKontextMultiReferenceLatentMethod → positive CONDITIONING
├─ TextEncodeQwenImageEditPlus (negative, empty) → FluxKontextMultiReferenceLatentMethod → negative CONDITIONING
└─ VAEEncode → LATENT
KSampler → VAEDecode → SaveImage
Official sampler settings:
| Variant | Steps | CFG | Sampler | Scheduler | Denoise | LoRA |
|---|
| Standard | 40 | 4.0 | euler | simple | 1.0 | none |
| Lightning | 4 | 1.0 | euler | simple | 1.0 | 2511-Lightning-4steps |
Note: The FluxKontextMultiReferenceLatentMethod and FluxKontextImageScale nodes may not be needed when using Comfy's official model files directly, but may be required with community-repackaged models.
XY Plot Technique (from Widgets.json)
For batch-testing multiple edit variations, use the Easy Nodes XY Plot system:
- Text Multiline nodes define parameter lists (e.g., directions, angles)
- Split String breaks them into indexed options
- easy textIndexSwitch selects one at a time
- easy promptReplace substitutes
{X}, {Y}, {Z} placeholders in the base prompt
- easy XYPlotAdvanced + easy XYInputs: PromptSR drives the sweep
- easy pipeIn bundles model/clip/vae/latent into a pipeline
This produces a grid image showing all combinations, useful for finding the best angle/distance/style for a given subject.
VRAM Considerations
- Qwen 2511 edit bf16: ~10GB VRAM
- CLIP (fp8): ~7GB VRAM
- VAE: ~200MB
- Total: ~17-18GB, fits comfortably on 24GB GPUs
- Always
clear_vram before loading if switching from another model family
Tips
- Upload source images first with
upload_image (action:"image") before building the workflow
- Match output resolution to the next pipeline step (e.g., 832x480 for WAN FLF) — but only on a generation graph. On an edit graph the size is the source's; resize the RESULT afterwards instead of sampling at the size you want
- Lightning LoRA + denoise 1.0 works well. The model handles structure preservation through conditioning
- Take an edit graph's
latent_image from a VAEEncode of the source, not from an EmptyLatentImage — at sub-1.0 denoise the empty latent decodes to a flat, near-uniform field (#2678), and at denoise 1.0 it is right only if its literal size happens to equal the geometry the encoder derived from the source, which is exactly the coincidence a VAEEncode removes (#2681). Neither failure errors. create_workflow (action:"validate") flags both pairings (partial_denoise_empty_latent, edit_reference_empty_latent)
- The lrzjason Pro variant is best for multi-image compositions where you need fine control over which images get VL-resized
- Use
get_workflow (action:"analyze") to understand any saved Qwen edit workflow before modifying or executing it. It returns a structured summary, not raw JSON. Only use get_workflow when you need the actual JSON for enqueue_workflow or create_workflow (action:"modify").
Sources
- Official: none found.
- Empirical: sampler values, wiring, and prompt notes from working graphs in
packs/ and observed renders; not a vendor prompting guide.