Files
aigen/docs/image-v2.md
T
TowstyandCursor 5fe91e6bae Show Flux and Krea as Image v2 engine cards, and run Krea on edit, compose, and refine.
Krea image-to-image uses encode plus denoise graphs, not Klein reference latents, and still fails closed if the Krea models are missing.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-29 07:29:31 -05:00

8.3 KiB
Raw Blame History

Image v1 vs Image v2

Image edit (v1) and Image v2 are siblings. v1 is unchanged. Do not mute-fix workflow_flux2_klein_edit.json (the dual-branch template that deletes 75:* or 92:* at runtime).

Image edit (v1) Image v2
Tab Image edit Image v2
Route POST /api/edit POST /api/v2/generate and POST /v2/generate
Graphs One file, prune unused branch Klein or Krea graphs per mode. Engine is always visible.
Second still Optional; switches branch Compose only; required. Mismatch is a 400
Mask — Refine only; required. Missing mask is a 400, not an Edit fallback
Input still Required Edit / Compose / Refine required. Generate ignores stills
LoRAs Optional user picker Consistency (Flux only). Concept LoRA on xAIGen only. Krea has no Consistency
Defaults 20 steps, CFG 1 Flux: 24 steps, CFG 4 (turbo: 8 / 1). Krea: 8 steps, CFG 1

Poll v2 jobs the same way as v1: GET /api/generate/{id}/stream and GET /api/generate/{id}.

Modes

  • edit — one image + text. Task is scene. Sending image_b is rejected. Use this to change lighting, background, or the whole frame from a single still.
  • compose — two images + text. image_b is required. No silent one-image fallback. Use this to lock a person from A and pull a face or outfit from B.
  • refine — one canvas + a painted mask + text. image_a and mask are required. image_b is ignored. Prompt only what should change inside the mask. Strength is denoise in that region (not CFG). Use this for face / hand / chest fixes without opening Comfy.
  • generate — text only. No image_a, image_b, or mask. Size is width × height (768 / 1024 / 1280), not megapixels from a still. Engine is Flux or Krea for every Image v2 mode. Video stays MiniMax / LTX.

Compose tasks:

  • identity — person from A; B reinforces the face; prompt changes pose/scene
  • outfit — person, body, pose, background from A; clothing only from B
  • face_lock — body/pose/scene from A; face from B
  • scene — A is the edit image; B is style/background

Refine and Generate do not prepend Edit/Compose role headers. The user prompt is the whole prompt.

Dummy tests:

  • Compose + two unrelated photos + “keep everything the same” must change the output. If it matches old v1 gens, still B is not connected.
  • Refine + last good gen + mask on the left breast + “red X painted on the left breast” at strength 0.4. Pass = X on the breast, face/pose/background stay. Fail = whole image regenerates, or nothing changes.
  • Generate + Flux + “a red cube on a white table, studio light” and no stills. Pass = a new image of that. Fail = missing-image error, or the output equals the last Edit/Compose still.
  • Generate + Krea + the same prompt. Pass = krea_v2_generate.json is queued. Fail = Flux graph runs under a Krea label, or a missing Krea model falls back to Flux.
  • Edit + Krea + one still. Pass = krea_v2_edit.json. Fail = Klein graph or Generate empty latent.

Nodes the mapper patches

Edit and Compose:

Node Role
1 Load Image A (inputs.image)
2 / 23 Scale to MP, lanczos (megapixels)
7 Concept LoRA (lora_name, strength_model, strength_clip)
8 Consistency LoRA (lora_name, strength_model, strength_clip)
9 Positive CLIPTextEncode.text (role header + user prompt on Compose)
10 Negative CLIPTextEncode.text
15 Seed
17 Steps (Flux2Scheduler; Klein-native, euler sampler)
18 CFG
21 SaveImage prefix

Compose only:

Node Role
22 Load Image B
23 Scale B
24 VAE encode B
25 / 26 Second ReferenceLatent (B fused into pos/neg)

Refine only (klein_v2_refine.json):

Node Role
1 Load canvas
30 Load mask (white = edit, black = keep)
2 Scale canvas to MP
31 / 32 / 34 Resize mask to canvas, take red channel, grow slightly
33 SetLatentNoiseMask on the encoded canvas
17 BasicScheduler.denoise — this is Strength, not CFG
19 Sampler starts from the masked canvas latent, not an empty Flux2 latent

Generate Flux (klein_v2_generate.json):

Node Role
14 EmptyFlux2LatentImage at request width × height
17 Flux2Scheduler steps + the same size
7 Concept LoRA (Klein file)
8 Consistency LoRA — removed at runtime when both strengths are 0
9 / 10 Prompt / negative from the API (no leftover widget text)

Generate Krea (krea_v2_generate.json):

Node Role
4 UNETLoader — krea2_turbo_mxfp8, then nvfp4, then fp8_scaled
5 CLIPLoader type krea2 — qwen3vl_4b_fp8_scaled.safetensors
6 VAELoader — qwen_image_vae.safetensors
14 EmptyLatentImage at request width × height
15 KSampler euler / simple, default 8 steps, CFG 1
9 / 10 Prompt / negative from the API

Krea Edit / Compose / Refine use the same UNET / CLIP / VAE. They encode the still (VAEEncode) and sample with denoise < 1. Krea has no Klein ReferenceLatent. Compose stitches A then B (ImageStitch) so both stills are in the latent. Missing Krea models fail the job. No Flux fallback.

Krea does not load Klein UNET, Klein CLIP, Klein VAE, Klein Consistency, or the Klein concept LoRA. Concept LoRA (klein_snofs_v1_4) is xAIGen only; AIGen bypasses that node. On Krea it stays off unless NUXT_KREA_CONCEPT_LORA is set on xAIGen.

There is no LoadImage on Generate. If the executed Generate graph has a required LoadImage, the job fails. v1 node 145 (muted T2I with a baked prompt) is not used.

If image_b is sent on Edit/Compose and the executed graph has fewer than two LoadImage nodes, the job fails. If Refine runs without a mask input or without denoise on the scheduler, the job fails.

Default sliders

Slider Default
Concept LoRA — Model xAIGen Flux: 0.65. Off on AIGen. 0 on Krea
Concept LoRA — CLIP xAIGen Flux: 0.35. Off on AIGen. 0 on Krea
Consistency — Model 0.70 on Flux Edit/Compose/Refine. Off on Generate
Consistency — CLIP 0.70 on Flux Edit/Compose/Refine. Off on Generate
Steps 24 (turbo 8)
CFG 4 (turbo 1)
Scale to MP 1.0 lanczos
Strength (denoise) 0.35 on Flux Refine. Krea Edit 0.75, Compose 0.70, Refine uses the slider. Hidden on Flux Edit/Compose/Generate. Range 0.15–0.75
Size Generate only. 1:1 1024×1024, 3:4 768×1024, 4:3 1024×768, 16:9 1280×768, 9:16 768×1280

When to use each Refine strength

Preset Strength SNOFS model / CLIP Consistency Use for
Face 0.25–0.35 (chip 0.28) 0.55 / 0.25 0.75 / 0.80 Glasses, cheeks, jaw. Keep the rest of the canvas.
Hand / chest 0.35–0.45 (chip 0.40) 0.65 / 0.30 0.70 / 0.70 Hands or chest, local contact. Dummy test uses 0.40.
Heavy 0.55 current sliders current sliders Stubborn region that barely moved at 0.40.

Do not raise Strength to rewrite the whole frame. If the face or background moves, the mask is too big or Strength is too high.

  • identity / face_lock — keep Consistency at 0.70 / 0.70 so the face from A (identity) or B (face_lock) holds. Do not drop Concept LoRA CLIP below ~0.30 or the skin/body read falls apart.
  • outfit — keep Concept LoRA Model ~0.65 so cloth reads; do not raise it past ~0.85 or it starts rewriting the body from A. Consistency stays at 0.70 so the person in A does not become the person in B.
  • scene / edit — start at the table defaults. Turbo is for drafts only.
  • refine — prompt only the masked change. Face vs hand/chest chips set Strength and LoRAs together.
  • generate / Flux — On xAIGen, Concept LoRA defaults 0.65 / 0.35. On AIGen the node is not loaded. Consistency stays 0. Turbo is 8 steps / CFG 1.
  • generate / Krea — 8 steps, CFG 1, no Consistency. Concept LoRA stays 0 unless a Krea concept file is configured.
  • edit / compose / refine / Krea — same 8 / CFG 1. Strength is denoise on the encoded still. Compose stitches A|B.

Presets on this tab store sliders + mode + task + engine (+ Strength when Refine, + size when Generate). They do not load v1 LoRA stacks or a cached latent. Missing engine on an old preset is Flux.