Add Image v2 Refine for on-device regional edits.

Mask painter and denoise strength sit beside Edit and Compose. The canned hand/chest prompt helper is gone so that text stays yours.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
Towsty
2026-08-28 22:03:20 -05:00
co-authored by Cursor
parent 21fe2412f9
commit a2895b8054
10 changed files with 825 additions and 56 deletions
+38 -9
View File
@@ -6,8 +6,9 @@ Image edit (v1) and Image v2 are siblings. v1 is unchanged. Do not mute-fix `wor
|---|---|---|
| Tab | Image edit | Image v2 |
| Route | `POST /api/edit` | `POST /api/v2/generate` and `POST /v2/generate` |
| Graphs | One file, prune unused branch | `klein_v2_edit.json` or `klein_v2_compose.json` |
| Graphs | One file, prune unused branch | `klein_v2_edit.json`, `klein_v2_compose.json`, or `klein_v2_refine.json` |
| Second still | Optional; switches branch | Compose only; required. Mismatch is a 400 |
| Mask | — | Refine only; required. Missing mask is a 400, not an Edit fallback |
| LoRAs | Optional user picker | SNOFS + Consistency, four strengths |
| Defaults | 20 steps, CFG 1 | 24 steps, CFG 4 (turbo: 8 / 1) |
@@ -15,8 +16,9 @@ Poll v2 jobs the same way as v1: `GET /api/generate/{id}/stream` and `GET /api/g
## Modes
- **edit** — one image + text. Task is `scene`. Sending `image_b` is rejected.
- **compose** — two images + text. `image_b` is required. No silent one-image fallback.
- **edit** — one image + text. Task is `scene`. Sending `image_b` is rejected. Use this to change lighting, background, or the whole frame from a single still.
- **compose** — two images + text. `image_b` is required. No silent one-image fallback. Use this to lock a person from A and pull a face or outfit from B.
- **refine** — one canvas + a painted mask + text. `image_a` and `mask` are required. `image_b` is ignored. Prompt only what should change inside the mask. Strength is denoise in that region (not CFG). Use this for face / hand / chest fixes without opening Comfy.
Compose tasks:
@@ -25,11 +27,16 @@ Compose tasks:
- **face_lock** — body/pose/scene from A; face from B
- **scene** — A is the edit image; B is style/background
Dummy test: Compose + two unrelated photos + “keep everything the same” must change the output. If it matches old v1 gens, still B is not connected.
Refine does not prepend Edit/Compose role headers.
Dummy tests:
- Compose + two unrelated photos + “keep everything the same” must change the output. If it matches old v1 gens, still B is not connected.
- Refine + last good gen + mask on the left breast + “red X painted on the left breast” at strength 0.4. Pass = X on the breast, face/pose/background stay. Fail = whole image regenerates, or nothing changes.
## Nodes the mapper patches
Both graphs:
Edit and Compose:
| Node | Role |
|------|------|
@@ -37,7 +44,7 @@ Both graphs:
| `2` / `23` | Scale to MP, lanczos (`megapixels`) |
| `7` | SNOFS LoRA (`lora_name`, `strength_model`, `strength_clip`) |
| `8` | Consistency LoRA (`lora_name`, `strength_model`, `strength_clip`) |
| `9` | Positive `CLIPTextEncode.text` (role header + user prompt) |
| `9` | Positive `CLIPTextEncode.text` (role header + user prompt on Compose) |
| `10` | Negative `CLIPTextEncode.text` |
| `15` | Seed |
| `17` | Steps (Flux2Scheduler; Klein-native, euler sampler) |
@@ -53,7 +60,19 @@ Compose only:
| `24` | VAE encode B |
| `25` / `26` | Second ReferenceLatent (B fused into pos/neg) |
If `image_b` is sent and the executed graph has fewer than two `LoadImage` nodes, the job fails.
Refine only (`klein_v2_refine.json`):
| Node | Role |
|------|------|
| `1` | Load canvas |
| `30` | Load mask (white = edit, black = keep) |
| `2` | Scale canvas to MP |
| `31` / `32` / `34` | Resize mask to canvas, take red channel, grow slightly |
| `33` | `SetLatentNoiseMask` on the encoded canvas |
| `17` | `BasicScheduler.denoise` — this is Strength, not CFG |
| `19` | Sampler starts from the masked canvas latent, not an empty Flux2 latent |
If `image_b` is sent on Edit/Compose and the executed graph has fewer than two `LoadImage` nodes, the job fails. If Refine runs without a mask input or without denoise on the scheduler, the job fails.
## Default sliders
@@ -66,13 +85,23 @@ If `image_b` is sent and the executed graph has fewer than two `LoadImage` nodes
| Steps | 24 (turbo 8) |
| CFG | 4 (turbo 1) |
| Scale to MP | 1.0 lanczos |
| Strength (denoise) | 0.35 on Refine only. Hidden on Edit/Compose. Range 0.15–0.75 |
There is no Denoise slider. This is reference-latent Klein, not inpaint.
### When to use each Refine strength
| Preset | Strength | SNOFS model / CLIP | Consistency | Use for |
|--------|----------|--------------------|-------------|---------|
| Face | 0.25–0.35 (chip 0.28) | 0.55 / 0.25 | 0.75 / 0.80 | Glasses, cheeks, jaw. Keep the rest of the canvas. |
| Hand / chest | 0.35–0.45 (chip 0.40) | 0.65 / 0.30 | 0.70 / 0.70 | Hands, breasts, local contact. Dummy test uses 0.40. |
| Heavy | 0.55 | current sliders | current sliders | Stubborn region that barely moved at 0.40. |
Do not raise Strength to rewrite the whole frame. If the face or background moves, the mask is too big or Strength is too high.
### NSFW identity vs outfit swap
- **identity / face_lock** — keep Consistency at 0.70 / 0.70 so the face from A (identity) or B (face_lock) holds. Do not drop SNOFS CLIP below ~0.30 or the skin/body read falls apart.
- **outfit** — keep SNOFS Model ~0.65 so cloth reads; do not raise it past ~0.85 or it starts rewriting the body from A. Consistency stays at 0.70 so the person in A does not become the person in B.
- **scene / edit** — start at the table defaults. Turbo is for drafts only.
- **refine** — prompt only the masked change. Face vs hand/chest chips set Strength and LoRAs together.
Presets on this tab store sliders + mode + task only. They do not load v1 LoRA stacks or a cached latent.
Presets on this tab store sliders + mode + task (+ Strength when Refine). They do not load v1 LoRA stacks or a cached latent.