Add Image v2 Generate for Klein text-to-image.

New sibling graph and mode so T2I does not borrow Edit, Compose, or Refine. Stills are ignored; size is width x height only.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
Towsty
2026-08-28 22:56:33 -05:00
co-authored by Cursor
parent a2895b8054
commit b35f9af10d
9 changed files with 511 additions and 112 deletions
+22 -5
View File
@@ -6,10 +6,11 @@ Image edit (v1) and Image v2 are siblings. v1 is unchanged. Do not mute-fix `wor
|---|---|---|
| Tab | Image edit | Image v2 |
| Route | `POST /api/edit` | `POST /api/v2/generate` and `POST /v2/generate` |
| Graphs | One file, prune unused branch | `klein_v2_edit.json`, `klein_v2_compose.json`, or `klein_v2_refine.json` |
| Graphs | One file, prune unused branch | `klein_v2_edit.json`, `klein_v2_compose.json`, `klein_v2_refine.json`, or `klein_v2_generate.json` |
| Second still | Optional; switches branch | Compose only; required. Mismatch is a 400 |
| Mask | — | Refine only; required. Missing mask is a 400, not an Edit fallback |
| LoRAs | Optional user picker | SNOFS + Consistency, four strengths |
| Input still | Required | Edit / Compose / Refine required. Generate ignores stills |
| LoRAs | Optional user picker | SNOFS + Consistency, four strengths. Generate defaults Consistency off |
| Defaults | 20 steps, CFG 1 | 24 steps, CFG 4 (turbo: 8 / 1) |
Poll v2 jobs the same way as v1: `GET /api/generate/{id}/stream` and `GET /api/generate/{id}`.
@@ -19,6 +20,7 @@ Poll v2 jobs the same way as v1: `GET /api/generate/{id}/stream` and `GET /api/g
- **edit** — one image + text. Task is `scene`. Sending `image_b` is rejected. Use this to change lighting, background, or the whole frame from a single still.
- **compose** — two images + text. `image_b` is required. No silent one-image fallback. Use this to lock a person from A and pull a face or outfit from B.
- **refine** — one canvas + a painted mask + text. `image_a` and `mask` are required. `image_b` is ignored. Prompt only what should change inside the mask. Strength is denoise in that region (not CFG). Use this for face / hand / chest fixes without opening Comfy.
- **generate** — text only. No `image_a`, `image_b`, or mask. Size is width × height (768 / 1024 / 1280), not megapixels from a still. Use this to make a new Klein still from a prompt.
Compose tasks:
@@ -27,12 +29,13 @@ Compose tasks:
- **face_lock** — body/pose/scene from A; face from B
- **scene** — A is the edit image; B is style/background
Refine does not prepend Edit/Compose role headers.
Refine and Generate do not prepend Edit/Compose role headers. The user prompt is the whole prompt.
Dummy tests:
- Compose + two unrelated photos + “keep everything the same” must change the output. If it matches old v1 gens, still B is not connected.
- Refine + last good gen + mask on the left breast + “red X painted on the left breast” at strength 0.4. Pass = X on the breast, face/pose/background stay. Fail = whole image regenerates, or nothing changes.
- Generate + “a red cube on a white table, studio light” and no stills. Pass = a new image of that. Fail = missing-image error, or the output equals the last Edit/Compose still.
## Nodes the mapper patches
@@ -72,6 +75,18 @@ Refine only (`klein_v2_refine.json`):
| `17` | `BasicScheduler.denoise` — this is Strength, not CFG |
| `19` | Sampler starts from the masked canvas latent, not an empty Flux2 latent |
Generate only (`klein_v2_generate.json`):
| Node | Role |
|------|------|
| `14` | `EmptyFlux2LatentImage` at request width × height |
| `17` | `Flux2Scheduler` steps + the same size |
| `7` | SNOFS LoRA |
| `8` | Consistency LoRA — removed at runtime when both strengths are 0 |
| `9` / `10` | Prompt / negative from the API (no leftover widget text) |
There is no `LoadImage`. If the executed Generate graph has a required LoadImage, the job fails. v1 node 145 (muted T2I with a baked prompt) is not used.
If `image_b` is sent on Edit/Compose and the executed graph has fewer than two `LoadImage` nodes, the job fails. If Refine runs without a mask input or without denoise on the scheduler, the job fails.
## Default sliders
@@ -85,7 +100,8 @@ If `image_b` is sent on Edit/Compose and the executed graph has fewer than two `
| Steps | 24 (turbo 8) |
| CFG | 4 (turbo 1) |
| Scale to MP | 1.0 lanczos |
| Strength (denoise) | 0.35 on Refine only. Hidden on Edit/Compose. Range 0.15–0.75 |
| Strength (denoise) | 0.35 on Refine only. Hidden on Edit/Compose/Generate. Range 0.15–0.75 |
| Size | Generate only. 1:1 1024×1024, 3:4 768×1024, 4:3 1024×768, 16:9 1280×768, 9:16 768×1280 |
### When to use each Refine strength
@@ -103,5 +119,6 @@ Do not raise Strength to rewrite the whole frame. If the face or background move
- **outfit** — keep SNOFS Model ~0.65 so cloth reads; do not raise it past ~0.85 or it starts rewriting the body from A. Consistency stays at 0.70 so the person in A does not become the person in B.
- **scene / edit** — start at the table defaults. Turbo is for drafts only.
- **refine** — prompt only the masked change. Face vs hand/chest chips set Strength and LoRAs together.
- **generate** — SFW scene = SNOFS 0 / 0. SNOF T2I = SNOFS 0.65 / 0.35. Consistency stays 0. Turbo is 8 steps / CFG 1.
Presets on this tab store sliders + mode + task (+ Strength when Refine). They do not load v1 LoRA stacks or a cached latent.
Presets on this tab store sliders + mode + task (+ Strength when Refine, + size when Generate). They do not load v1 LoRA stacks or a cached latent.