Fix Qwen 2.1 GGUF load: promote Q8 norms to F32 and wire TextEncode latent.

abenzerps Q8_0 ships 1D RMSNorms as packed Q8 (136 vs 128), which breaks
Comfy rms_rope; tagger now dequantizes small tensors and the graph uses
TextEncodeQwenImage21's 64-ch latent plus AuraFlow shift.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
Towsty
2026-09-20 16:58:00 -05:00
co-authored by Cursor
parent 68731916ca
commit 608f7d8d47
5 changed files with 200 additions and 39 deletions
+9 -1
View File
@@ -21,6 +21,14 @@ Custom node: `custom_nodes\ComfyUI-GGUF` (`UnetLoaderGGUF`). Do not install a se
powershell -ExecutionPolicy Bypass -File scripts\setup-qwen21.ps1
```
DiT download uses the exact host invocation:
```powershell
hf download hf://abenzerps/Qwen-Image-2.1-Uncensored-GGUF/qwen-image-2.1-Q8_0.gguf
```
That lands in the Hugging Face hub cache; the setup script copies it into Shared `diffusion_models`. The abenzerps GGUF ships with `kv_count=0` (no `general.architecture`) **and** Q8_0-quantized 1D RMSNorm weights (logical 128 → packed 136), which breaks Comfy’s fused `rms_rope`. `scripts/tag-qwen21-gguf.py` rewrites the file with `general.architecture=qwen_image` and promotes small/1D tensors to F32 so city96 `UnetLoaderGGUF` can load it.
Restart Comfy **only when idle** (`COMFY_CONTROL_URL/status` → `gpu.busy=false`). Confirm `object_info` lists `UnetLoaderGGUF` and the three filenames.
## App
@@ -28,7 +36,7 @@ Restart Comfy **only when idle** (`COMFY_CONTROL_URL/status` → `gpu.busy=false
- Engine key: `qwen21` · UI label: **Qwen 2.1**
- Generate (T2I) only in this build. Edit / Compose / Iterate / Video disable with: “Qwen 2.1 is T2I in this build”.
- Sampler defaults: euler / simple / cfg **1** / steps **25** · `ModelSamplingAuraFlow` shift **3.1**
- Default canvas **1024×1024** (aspect 16:9 / 9:16 / 1:1 honored, multiples of 32). Drop to 768 if VRAM errors.
- Default canvas **1024×1024** (aspect 16:9 / 9:16 / 1:1 → long-edge square via `TextEncodeQwenImage21` resolution; multiples of 32). Drop to 768 if VRAM errors.
- No hero / locks / Klein LoRA stack on this engine.
Graph: `server/assets/studio2_qwen21_t2i.json`.