Files
aigen/docs/qwen21.md
T
TowstyandCursor 50319e03a9 Map Photos roles to Qwen image_1/image_2 and inject encode tags silently.
Mention phrases become <imageN> only in the TextEncode string; the inspector never shows Comfy tokens.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-26 16:01:49 -05:00

53 lines
5.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Qwen Image 2.1 (engine `qwen21`) — host weights + Comfy graph
## Host paths (this RTX 5080 box)
| Role | Resolved path |
|---|---|
| ComfyUI root (Klein / host agent `ComfyUI (1)`) | `C:\Users\ianjm\AppData\Local\Comfy-Desktop\ComfyUI-Installs\ComfyUI (1)\ComfyUI` |
| Models root (Desktop Shared via `shared_model_paths.yaml`) | `C:\Users\ianjm\AppData\Local\Comfy-Desktop\ComfyUI-Shared\models` |
Weights must land under Shared so every Desktop instance sees them:
- `diffusion_models\qwen-image-2.1-Q8_0.gguf` (~7.59 GiB) — from `abenzerps/Qwen-Image-2.1-Uncensored-GGUF`
- `text_encoders\qwen3vl_8b_int8_convrot.safetensors` — from `Comfy-Org/Qwen-Image-2.1` (INT8; BF16 8B VL OOMs on 16 GB)
- `vae\qwen_image_2.1_vae_bf16.safetensors` — from `Comfy-Org/Qwen-Image-2.1` (**not** the old Qwen-Image 1.0 VAE)
Custom node: `custom_nodes\ComfyUI-GGUF` (`UnetLoaderGGUF`). Do not install a second GGUF pack.
## Setup
```powershell
powershell -ExecutionPolicy Bypass -File scripts\setup-qwen21.ps1
```
DiT download uses the exact host invocation:
```powershell
hf download hf://abenzerps/Qwen-Image-2.1-Uncensored-GGUF/qwen-image-2.1-Q8_0.gguf
```
That lands in the Hugging Face hub cache; the setup script copies it into Shared `diffusion_models`. The abenzerps GGUF ships with `kv_count=0` (no `general.architecture`) **and** Q8_0-quantized 1D RMSNorm weights (logical 128 → packed 136), which breaks Comfy’s fused `rms_rope`. `scripts/tag-qwen21-gguf.py` rewrites the file with `general.architecture=qwen_image` and promotes small/1D tensors to F32 so city96 `UnetLoaderGGUF` can load it.
Restart Comfy **only when idle** (`COMFY_CONTROL_URL/status` → `gpu.busy=false`). Confirm `object_info` lists `UnetLoaderGGUF` and the three filenames.
## App
- Engine key: `qwen21` · UI label: **Qwen 2.1**
- **Generate** (T2I) and **Edit** (same checkpoint, second graph). Compose / Iterate / Video / Extend / Music stay disabled unless a graph exists.
- Edit maps Photos roles → graph sockets: Photo to change → `images.image_1`, Outfit / object or Extra → `image_2`. The inspector never shows `<image1>`; the runner injects tags into the TextEncode string on submit (and expands Mention phrases like “this photo” / “the outfit photo”). If `<image1>` is still missing, prepend the keep-identity stanza. Klein hero-ref is off for Qwen.
- Face / outfit locks stay Klein semantics — they do not drive Qwen slots.
- Sampler defaults: euler / simple / cfg **1** / steps **25**. Edit uses `QwenImage21Cache` (device auto, dtype int8). T2I uses `ModelSamplingAuraFlow` shift **3.1**.
- Default Generate canvas follows the bench Aspect control on the Qwen-safe ~1 MP grid (`EmptyLatentImage`): 1:1 → 1024×1024, 16:9 → 1536×864, 9:16 → 864×1536 (and the other table rows). PE `wh_ratio` is advisory only and never sizes the canvas. Do not use native 2K bins on 16 GB.
- Edit follows `image_1` via the encode node’s latent (resolution long-edge ~1024). Do not inject a picker EmptyLatentImage onto the edit sampler.
- Optional **Enhance prompt** (off by default): runs a separate PE-only Comfy graph (`CLIPLoader` + rewrite node), then frees VRAM and queues the existing T2I/Edit graph with the rewritten prompt. Never loads PE CLIP + DiT together on 16 GB. Fail closed if `parse_ok` is false or the rewrite is empty.
- Generate + Enhance → PE-T2I only (`pe_t2i`). No images.
- Edit + Enhance → PE-I2I only (`pe_i2i`) with the same Start still on `image_1`. After PE, the sample prompt is always the keep-identity stanza + typed instruction first (`<image1>` in front). A PE rewrite is appended only when it is an edit directive; T2I-style observer captions (`The image is…`) are dropped (`enhance.skippedAsDescribe`) and the stanza + raw remain. Never let the PE chunk be the entire prompt. Library stores `prompt` (TextEncode string) and `promptRaw` (Typed). Never fall back to PE-T2I on Edit.
- If PE refuses or returns an empty/gutted rewrite, the job continues with `promptRaw` (Edit still gets the keep-identity stanza). Details shows `Enhance skipped (model refused) — used your prompt.` (`enhance.refused`).
- Host PE system prompts live in `host/qwen21-pe-prompts/` and are copied onto the node pack by `scripts/setup-qwen21.ps1` (official steps + Adult appendix). Edit prompt must lead with an operation and `<image1>` — never “The image is a photograph of…”. Do not swap Heretic PE weights on 16 GB.
- PE weights (int8 only): `text_encoders\qwen3.5_9b_qwen_image_2.1_pe_{t2i,i2i}.int8_convrot.safetensors` — do not replace the image TE `qwen3vl_8b_int8_convrot.safetensors`.
- Optional **Turbo** (off by default): Viggle DMD LoRA `Qwen-Image-2.1-viggle-turbo-v0.2.1-6step-lora-r128.safetensors` on the same Q8 GGUF via `ViggleTurboLora` + `ViggleTurboSigmas` (`1.0, 0.9375, 0.875, 0.75, 0.5, 0.25`), 6 steps, CFG 1, empty negative. Not a new engine. Enhance prompt stays compatible and is recommended with Turbo.
- Custom node: `ComfyUI-Viggle-Turbo` (`viggle_turbo.py`). Do not merge the LoRA with stock `LoraLoaderModelOnly` (lossy on int8/bf16).
Graphs: `server/assets/studio2_qwen21_t2i.json`, `server/assets/studio2_qwen21_edit.json`, `server/assets/studio2_qwen21_t2i_turbo.json`, `server/assets/studio2_qwen21_edit_turbo.json`, `server/assets/studio2_qwen21_pe_t2i.json`, `server/assets/studio2_qwen21_pe_edit.json`.