Files
aigen/docs/image-v2.md
T
TowstyandCursor f6663444f5 Add Krea as a second Image v2 Generate engine beside Flux.
Generate can queue krea_v2_generate.json without sharing Klein encoders or falling back if Krea assets are missing.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-29 07:05:01 -05:00

138 lines
7.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Image v1 vs Image v2
Image edit (v1) and Image v2 are siblings. v1 is unchanged. Do not mute-fix `workflow_flux2_klein_edit.json` (the dual-branch template that deletes `75:*` or `92:*` at runtime).
| | Image edit (v1) | Image v2 |
|---|---|---|
| Tab | Image edit | Image v2 |
| Route | `POST /api/edit` | `POST /api/v2/generate` and `POST /v2/generate` |
| Graphs | One file, prune unused branch | `klein_v2_edit.json`, `klein_v2_compose.json`, `klein_v2_refine.json`, `klein_v2_generate.json`, or `krea_v2_generate.json` |
| Second still | Optional; switches branch | Compose only; required. Mismatch is a 400 |
| Mask | — | Refine only; required. Missing mask is a 400, not an Edit fallback |
| Input still | Required | Edit / Compose / Refine required. Generate ignores stills |
| LoRAs | Optional user picker | Concept LoRA + Consistency (Flux). Generate Consistency off. Krea Generate has no Consistency |
| Defaults | 20 steps, CFG 1 | Flux: 24 steps, CFG 4 (turbo: 8 / 1). Krea Generate: 8 steps, CFG 1 |
Poll v2 jobs the same way as v1: `GET /api/generate/{id}/stream` and `GET /api/generate/{id}`.
## Modes
- **edit** — one image + text. Task is `scene`. Sending `image_b` is rejected. Use this to change lighting, background, or the whole frame from a single still.
- **compose** — two images + text. `image_b` is required. No silent one-image fallback. Use this to lock a person from A and pull a face or outfit from B.
- **refine** — one canvas + a painted mask + text. `image_a` and `mask` are required. `image_b` is ignored. Prompt only what should change inside the mask. Strength is denoise in that region (not CFG). Use this for face / hand / chest fixes without opening Comfy.
- **generate** — text only. No `image_a`, `image_b`, or mask. Size is width × height (768 / 1024 / 1280), not megapixels from a still. Engine is Flux (`klein_v2_generate.json`) or Krea (`krea_v2_generate.json`). Edit / Compose / Refine stay Flux.
Compose tasks:
- **identity** — person from A; B reinforces the face; prompt changes pose/scene
- **outfit** — person, body, pose, background from A; clothing only from B
- **face_lock** — body/pose/scene from A; face from B
- **scene** — A is the edit image; B is style/background
Refine and Generate do not prepend Edit/Compose role headers. The user prompt is the whole prompt.
Dummy tests:
- Compose + two unrelated photos + “keep everything the same” must change the output. If it matches old v1 gens, still B is not connected.
- Refine + last good gen + mask on the left breast + “red X painted on the left breast” at strength 0.4. Pass = X on the breast, face/pose/background stay. Fail = whole image regenerates, or nothing changes.
- Generate + Flux + “a red cube on a white table, studio light” and no stills. Pass = a new image of that. Fail = missing-image error, or the output equals the last Edit/Compose still.
- Generate + Krea + the same prompt. Pass = `krea_v2_generate.json` is queued. Fail = Flux graph runs under a Krea label, or a missing Krea model falls back to Flux.
## Nodes the mapper patches
Edit and Compose:
| Node | Role |
|------|------|
| `1` | Load Image A (`inputs.image`) |
| `2` / `23` | Scale to MP, lanczos (`megapixels`) |
| `7` | Concept LoRA (`lora_name`, `strength_model`, `strength_clip`) |
| `8` | Consistency LoRA (`lora_name`, `strength_model`, `strength_clip`) |
| `9` | Positive `CLIPTextEncode.text` (role header + user prompt on Compose) |
| `10` | Negative `CLIPTextEncode.text` |
| `15` | Seed |
| `17` | Steps (Flux2Scheduler; Klein-native, euler sampler) |
| `18` | CFG |
| `21` | SaveImage prefix |
Compose only:
| Node | Role |
|------|------|
| `22` | Load Image B |
| `23` | Scale B |
| `24` | VAE encode B |
| `25` / `26` | Second ReferenceLatent (B fused into pos/neg) |
Refine only (`klein_v2_refine.json`):
| Node | Role |
|------|------|
| `1` | Load canvas |
| `30` | Load mask (white = edit, black = keep) |
| `2` | Scale canvas to MP |
| `31` / `32` / `34` | Resize mask to canvas, take red channel, grow slightly |
| `33` | `SetLatentNoiseMask` on the encoded canvas |
| `17` | `BasicScheduler.denoise` — this is Strength, not CFG |
| `19` | Sampler starts from the masked canvas latent, not an empty Flux2 latent |
Generate Flux (`klein_v2_generate.json`):
| Node | Role |
|------|------|
| `14` | `EmptyFlux2LatentImage` at request width × height |
| `17` | `Flux2Scheduler` steps + the same size |
| `7` | Concept LoRA (Klein file) |
| `8` | Consistency LoRA — removed at runtime when both strengths are 0 |
| `9` / `10` | Prompt / negative from the API (no leftover widget text) |
Generate Krea (`krea_v2_generate.json`):
| Node | Role |
|------|------|
| `4` | `UNETLoader` — `krea2_turbo_mxfp8`, then `nvfp4`, then `fp8_scaled` |
| `5` | `CLIPLoader` type `krea2` — `qwen3vl_4b_fp8_scaled.safetensors` |
| `6` | `VAELoader` — `qwen_image_vae.safetensors` |
| `14` | `EmptyLatentImage` at request width × height |
| `15` | `KSampler` euler / simple, default 8 steps, CFG 1 |
| `9` / `10` | Prompt / negative from the API |
Krea does not load Klein UNET, Klein CLIP, Klein VAE, Klein Consistency, or the Klein concept LoRA. Missing Krea diffusion / text encoder / VAE fails the job. Concept LoRA on Krea is off unless `NUXT_KREA_CONCEPT_LORA` is set.
There is no `LoadImage`. If the executed Generate graph has a required LoadImage, the job fails. v1 node 145 (muted T2I with a baked prompt) is not used.
If `image_b` is sent on Edit/Compose and the executed graph has fewer than two `LoadImage` nodes, the job fails. If Refine runs without a mask input or without denoise on the scheduler, the job fails.
## Default sliders
| Slider | Default |
|--------|---------|
| Concept LoRA — Model | 0.65 on Flux. 0 on Krea |
| Concept LoRA — CLIP | 0.35 on Flux. 0 on Krea |
| Consistency — Model | 0.70 on Flux Edit/Compose/Refine. Off on Generate |
| Consistency — CLIP | 0.70 on Flux Edit/Compose/Refine. Off on Generate |
| Steps | 24 (turbo 8) |
| CFG | 4 (turbo 1) |
| Scale to MP | 1.0 lanczos |
| Strength (denoise) | 0.35 on Refine only. Hidden on Edit/Compose/Generate. Range 0.15–0.75 |
| Size | Generate only. 1:1 1024×1024, 3:4 768×1024, 4:3 1024×768, 16:9 1280×768, 9:16 768×1280 |
### When to use each Refine strength
| Preset | Strength | SNOFS model / CLIP | Consistency | Use for |
|--------|----------|--------------------|-------------|---------|
| Face | 0.25–0.35 (chip 0.28) | 0.55 / 0.25 | 0.75 / 0.80 | Glasses, cheeks, jaw. Keep the rest of the canvas. |
| Hand / chest | 0.35–0.45 (chip 0.40) | 0.65 / 0.30 | 0.70 / 0.70 | Hands or chest, local contact. Dummy test uses 0.40. |
| Heavy | 0.55 | current sliders | current sliders | Stubborn region that barely moved at 0.40. |
Do not raise Strength to rewrite the whole frame. If the face or background moves, the mask is too big or Strength is too high.
- **identity / face_lock** — keep Consistency at 0.70 / 0.70 so the face from A (identity) or B (face_lock) holds. Do not drop Concept LoRA CLIP below ~0.30 or the skin/body read falls apart.
- **outfit** — keep Concept LoRA Model ~0.65 so cloth reads; do not raise it past ~0.85 or it starts rewriting the body from A. Consistency stays at 0.70 so the person in A does not become the person in B.
- **scene / edit** — start at the table defaults. Turbo is for drafts only.
- **refine** — prompt only the masked change. Face vs hand/chest chips set Strength and LoRAs together.
- **generate / Flux** — Concept LoRA defaults 0.65 / 0.35. Consistency stays 0. Turbo is 8 steps / CFG 1.
- **generate / Krea** — 8 steps, CFG 1, no Consistency. Concept LoRA stays 0 unless a Krea concept file is configured.
Presets on this tab store sliders + mode + task + engine (+ Strength when Refine, + size when Generate). They do not load v1 LoRA stacks or a cached latent. Missing engine on an old preset is Flux.