diff --git a/docs/image-v2.md b/docs/image-v2.md index 7c4e844..2572d90 100644 --- a/docs/image-v2.md +++ b/docs/image-v2.md @@ -6,10 +6,11 @@ Image edit (v1) and Image v2 are siblings. v1 is unchanged. Do not mute-fix `wor |---|---|---| | Tab | Image edit | Image v2 | | Route | `POST /api/edit` | `POST /api/v2/generate` and `POST /v2/generate` | -| Graphs | One file, prune unused branch | `klein_v2_edit.json`, `klein_v2_compose.json`, or `klein_v2_refine.json` | +| Graphs | One file, prune unused branch | `klein_v2_edit.json`, `klein_v2_compose.json`, `klein_v2_refine.json`, or `klein_v2_generate.json` | | Second still | Optional; switches branch | Compose only; required. Mismatch is a 400 | | Mask | — | Refine only; required. Missing mask is a 400, not an Edit fallback | -| LoRAs | Optional user picker | SNOFS + Consistency, four strengths | +| Input still | Required | Edit / Compose / Refine required. Generate ignores stills | +| LoRAs | Optional user picker | SNOFS + Consistency, four strengths. Generate defaults Consistency off | | Defaults | 20 steps, CFG 1 | 24 steps, CFG 4 (turbo: 8 / 1) | Poll v2 jobs the same way as v1: `GET /api/generate/{id}/stream` and `GET /api/generate/{id}`. @@ -19,6 +20,7 @@ Poll v2 jobs the same way as v1: `GET /api/generate/{id}/stream` and `GET /api/g - **edit** — one image + text. Task is `scene`. Sending `image_b` is rejected. Use this to change lighting, background, or the whole frame from a single still. - **compose** — two images + text. `image_b` is required. No silent one-image fallback. Use this to lock a person from A and pull a face or outfit from B. - **refine** — one canvas + a painted mask + text. `image_a` and `mask` are required. `image_b` is ignored. Prompt only what should change inside the mask. Strength is denoise in that region (not CFG). Use this for face / hand / chest fixes without opening Comfy. +- **generate** — text only. No `image_a`, `image_b`, or mask. Size is width × height (768 / 1024 / 1280), not megapixels from a still. Use this to make a new Klein still from a prompt. Compose tasks: @@ -27,12 +29,13 @@ Compose tasks: - **face_lock** — body/pose/scene from A; face from B - **scene** — A is the edit image; B is style/background -Refine does not prepend Edit/Compose role headers. +Refine and Generate do not prepend Edit/Compose role headers. The user prompt is the whole prompt. Dummy tests: - Compose + two unrelated photos + “keep everything the same” must change the output. If it matches old v1 gens, still B is not connected. - Refine + last good gen + mask on the left breast + “red X painted on the left breast” at strength 0.4. Pass = X on the breast, face/pose/background stay. Fail = whole image regenerates, or nothing changes. +- Generate + “a red cube on a white table, studio light” and no stills. Pass = a new image of that. Fail = missing-image error, or the output equals the last Edit/Compose still. ## Nodes the mapper patches @@ -72,6 +75,18 @@ Refine only (`klein_v2_refine.json`): | `17` | `BasicScheduler.denoise` — this is Strength, not CFG | | `19` | Sampler starts from the masked canvas latent, not an empty Flux2 latent | +Generate only (`klein_v2_generate.json`): + +| Node | Role | +|------|------| +| `14` | `EmptyFlux2LatentImage` at request width × height | +| `17` | `Flux2Scheduler` steps + the same size | +| `7` | SNOFS LoRA | +| `8` | Consistency LoRA — removed at runtime when both strengths are 0 | +| `9` / `10` | Prompt / negative from the API (no leftover widget text) | + +There is no `LoadImage`. If the executed Generate graph has a required LoadImage, the job fails. v1 node 145 (muted T2I with a baked prompt) is not used. + If `image_b` is sent on Edit/Compose and the executed graph has fewer than two `LoadImage` nodes, the job fails. If Refine runs without a mask input or without denoise on the scheduler, the job fails. ## Default sliders @@ -85,7 +100,8 @@ If `image_b` is sent on Edit/Compose and the executed graph has fewer than two ` | Steps | 24 (turbo 8) | | CFG | 4 (turbo 1) | | Scale to MP | 1.0 lanczos | -| Strength (denoise) | 0.35 on Refine only. Hidden on Edit/Compose. Range 0.15–0.75 | +| Strength (denoise) | 0.35 on Refine only. Hidden on Edit/Compose/Generate. Range 0.15–0.75 | +| Size | Generate only. 1:1 1024×1024, 3:4 768×1024, 4:3 1024×768, 16:9 1280×768, 9:16 768×1280 | ### When to use each Refine strength @@ -103,5 +119,6 @@ Do not raise Strength to rewrite the whole frame. If the face or background move - **outfit** — keep SNOFS Model ~0.65 so cloth reads; do not raise it past ~0.85 or it starts rewriting the body from A. Consistency stays at 0.70 so the person in A does not become the person in B. - **scene / edit** — start at the table defaults. Turbo is for drafts only. - **refine** — prompt only the masked change. Face vs hand/chest chips set Strength and LoRAs together. +- **generate** — SFW scene = SNOFS 0 / 0. SNOF T2I = SNOFS 0.65 / 0.35. Consistency stays 0. Turbo is 8 steps / CFG 1. -Presets on this tab store sliders + mode + task (+ Strength when Refine). They do not load v1 LoRA stacks or a cached latent. +Presets on this tab store sliders + mode + task (+ Strength when Refine, + size when Generate). They do not load v1 LoRA stacks or a cached latent. diff --git a/pages/index.vue b/pages/index.vue index 2d91d7e..c63bb25 100644 --- a/pages/index.vue +++ b/pages/index.vue @@ -65,7 +65,9 @@

Input

{{ studioMode === 'editv2' - ? (v2Mode === 'refine' + ? (v2Mode === 'generate' + ? 'Generate: no reference image. Prompt only. Klein 9B Base text-to-image.' + : v2Mode === 'refine' ? 'Refine: paint a mask on Still A. Prompt only what should change in the painted area. Strength is denoise in that region.' : v2Mode === 'compose' ? 'Compose: still A is the person/body. Still B is the face or outfit. Two stills required — no silent one-image fallback.' @@ -76,6 +78,7 @@

Refine + -