Keep last-frame continuity when character stills are attached.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
Towsty
2026-08-30 08:21:57 -05:00
co-authored by Cursor
parent 5d828afbd7
commit 0fa64ea765
10 changed files with 29 additions and 51 deletions
+1 -1
View File
@@ -100,7 +100,7 @@ Generate Krea (`krea_v2_generate.json`):
Krea Edit / Compose / Refine use the same UNET / CLIP / VAE. They encode the still (`VAEEncode`) and sample with denoise &lt; 1. Krea has no Klein `ReferenceLatent`. Compose stitches A then B (`ImageStitch`) so both stills are in the latent. Missing Krea models fail the job. No Flux fallback. Krea Edit / Compose / Refine use the same UNET / CLIP / VAE. They encode the still (`VAEEncode`) and sample with denoise &lt; 1. Krea has no Klein `ReferenceLatent`. Compose stitches A then B (`ImageStitch`) so both stills are in the latent. Missing Krea models fail the job. No Flux fallback.
Krea does not load Klein UNET, Klein CLIP, Klein VAE, Klein Consistency, or the Klein concept LoRA. Concept LoRA (`klein_snofs_v1_4`) is xAIGen only; AIGen bypasses that node. On Krea it stays off unless `NUXT_KREA_CONCEPT_LORA` is set on xAIGen. Krea does not load Klein UNET, Klein CLIP, Klein VAE, Klein Consistency, or the Klein concept LoRA. Concept LoRA (`xaigen-klein_snofs_v1_4`) is xAIGen only; AIGen bypasses that node. On Krea it stays off unless `NUXT_KREA_CONCEPT_LORA` is set on xAIGen.
There is no `LoadImage` on Generate. If the executed Generate graph has a required LoadImage, the job fails. v1 node 145 (muted T2I with a baked prompt) is not used. There is no `LoadImage` on Generate. If the executed Generate graph has a required LoadImage, the job fails. v1 node 145 (muted T2I with a baked prompt) is not used.
+14 -34
View File
@@ -273,10 +273,10 @@
/> />
<p v-if="shotScriptMode && studioMode === 'video'" class="mb-2 text-xs text-zinc-500"> <p v-if="shotScriptMode && studioMode === 'video'" class="mb-2 text-xs text-zinc-500">
Paste a block, or load a saved Director script. Text before the first marker is shot 1. Put <span class="font-medium text-zinc-300">shot 2</span>, <span class="font-medium text-zinc-300">shot 3</span> on their own lines to queue extensions — a line like <span class="font-medium text-zinc-300">Shot Type</span> will not start a new shot. Every shot uses the duration slider. Global locks are inherited; each shot should add one beat. Paste a block, or load a saved Director script. Text before the first marker is shot 1. Put <span class="font-medium text-zinc-300">shot 2</span>, <span class="font-medium text-zinc-300">shot 3</span> on their own lines to queue extensions — a line like <span class="font-medium text-zinc-300">Shot Type</span> will not start a new shot. Every shot uses the duration slider. Global locks are inherited; each shot should add one beat.
<span v-if="identityPrompting"> The Comfy identity prefix is prepended to every shot, and the same Picture 1 / extra stills are sent with each shot.</span> <span v-if="identityPrompting"> Character stills ride along as extra pictures. Picture 1 is the current frame (last frame after shot 1), not the original start still.</span>
</p> </p>
<p v-else-if="identityPrompting" class="mb-2 text-xs text-zinc-500"> <p v-else-if="identityPrompting" class="mb-2 text-xs text-zinc-500">
Write camera, action, and audio. That is the dynamic action prompt. We concatenate the <span class="font-medium text-zinc-300">&lt;Picture 1&gt;</span> identity prefix in front of it and send the result to <span class="font-medium text-zinc-300">Reference to Video</span>, not Image to Video. Write camera, action, and audio. Picture 1 is the current frame. Extra stills lock face and body only — they do not reset clothes. MiniMax still cannot hard-lock a last frame and character refs in one node, so this path is Reference to Video with the last frame as Picture 1.
</p> </p>
<div <div
class="grid gap-3" class="grid gap-3"
@@ -667,25 +667,18 @@
:class="minimaxGraph === 'v2' ? 'border-amber-300 bg-amber-400/10 text-amber-100' : 'border-white/10 hover:border-white/20'" :class="minimaxGraph === 'v2' ? 'border-amber-300 bg-amber-400/10 text-amber-100' : 'border-white/10 hover:border-white/20'"
@click="minimaxGraph = 'v2'" @click="minimaxGraph = 'v2'"
> >
<span class="block font-semibold">v2 · identity refs</span> <span class="block font-semibold">v2 · last frame + stills</span>
<span class="text-xs text-zinc-400">Optional lock with up to 4 stills</span> <span class="text-xs text-zinc-400">Optional character stills. Last frame stays the start.</span>
</button> </button>
</div> </div>
</div> </div>
<div v-if="videoEngine === 'minimax' && videoStart === 'still' && minimaxGraph === 'v2'" class="rounded-2xl border border-white/10 bg-zinc-950/40 px-4 py-3 space-y-3"> <div v-if="videoEngine === 'minimax' && videoStart === 'still' && minimaxGraph === 'v2'" class="rounded-2xl border border-white/10 bg-zinc-950/40 px-4 py-3 space-y-3">
<label class="flex cursor-pointer items-center justify-between gap-3 text-sm" :class="extraPersonLock ? 'cursor-not-allowed opacity-60' : ''"> <div>
<span> <p class="text-sm font-medium text-zinc-200">Character stills</p>
<span class="block font-medium text-zinc-200">Use identity references</span> <p class="mt-0.5 text-xs text-zinc-500">Optional. Extra views of the same person — face, hair, body. Shot 2+ still starts from the last frame (clothes and pose stay). Empty slots are left out. A named extra person turns these off.</p>
<span class="mt-0.5 block text-xs text-zinc-500">Off uses last-frame I2V after shot 1. On keeps the start still as Picture 1 — extra stills are more views of that same person, never a second character. A named extra person forces last-frame I2V.</span> </div>
</span> <div v-if="!extraPersonLock" class="grid grid-cols-2 gap-2 sm:grid-cols-4">
<span class="relative inline-flex h-6 w-11 shrink-0 items-center">
<input v-model="useIdentityRefs" type="checkbox" class="peer sr-only" role="switch" :disabled="extraPersonLock">
<span class="absolute inset-0 rounded-full bg-zinc-700 transition peer-checked:bg-amber-400" />
<span class="absolute left-0.5 top-0.5 h-5 w-5 rounded-full bg-white transition peer-checked:translate-x-5" />
</span>
</label>
<div v-if="useIdentityRefs && !extraPersonLock" class="grid grid-cols-2 gap-2 sm:grid-cols-4">
<div <div
v-for="slot in 4" v-for="slot in 4"
:key="slot" :key="slot"
@@ -723,8 +716,8 @@
</button> </button>
</div> </div>
</div> </div>
<p v-if="extraPersonLock" class="text-xs text-zinc-500">Identity refs stay off while a second named person is attached. MiniMax will not see the extra stills; shot 2+ starts from the parent clip’s last frame.</p> <p v-if="extraPersonLock" class="text-xs text-zinc-500">Character stills stay off while a second named person is attached. Shot 2+ starts from the last frame.</p>
<p v-else-if="useIdentityRefs" class="text-xs text-zinc-500">Picture 1 is the start still and stays Picture 1 on every shot, including queued extensions. Extra stills map to Picture 2–5 only if you add them — extra views of that same person, not a second character or a vehicle. Empty slots are left out of the graph, not filled with copies of Picture 1.</p> <p v-else class="text-xs text-zinc-500">Picture 1 is the start still on shot 1 and the last frame after that. Extra stills are Picture 2–5. They do not put the start-still outfit back on.</p>
</div> </div>
</div> </div>
@@ -2449,7 +2442,6 @@ const imageV2LoraOptions = computed(() => mergeLoraOptions(
filterLorasForImageEngine(imageLoras.value, v2Engine.value), filterLorasForImageEngine(imageLoras.value, v2Engine.value),
filterLorasForImageEngine(imageV2LoraStack.value.map(item => item.name), v2Engine.value) filterLorasForImageEngine(imageV2LoraStack.value.map(item => item.name), v2Engine.value)
)) ))
const useIdentityRefs = ref(false)
const identityRefs = ref<(File | null)[]>([null, null, null, null]) const identityRefs = ref<(File | null)[]>([null, null, null, null])
const identityPreviews = ref(['', '', '', '']) const identityPreviews = ref(['', '', '', ''])
const globalLocks = ref('') const globalLocks = ref('')
@@ -2474,12 +2466,13 @@ const extraPersonLock = computed(() => shouldForceLastFrameI2V(
mergePermanenceRefs(familyPermanenceRefs.value, Object.values(shotPermanenceRefs.value).flat()), mergePermanenceRefs(familyPermanenceRefs.value, Object.values(shotPermanenceRefs.value).flat()),
globalLocks.value globalLocks.value
)) ))
const useIdentityRefs = computed(() => !extraPersonLock.value && identityRefs.value.some(Boolean))
const permanenceComfyNote = computed(() => { const permanenceComfyNote = computed(() => {
if (extraPersonLock.value) { if (extraPersonLock.value) {
return 'A second named person keeps last-frame I2V on. MiniMax never sees the man/bench photos; those names only go in the inherit line.' return 'A second named person keeps last-frame I2V on. MiniMax never sees the man/bench photos; those names only go in the inherit line.'
} }
if (studioMode.value === 'video' && videoWorkflow.value === 'v2' && useIdentityRefs.value) { if (studioMode.value === 'video' && videoWorkflow.value === 'v2' && useIdentityRefs.value) {
return 'Identity-ref still only maps extra views of the start-still person, never a second character. Permanence stills name the lock text; they are not Picture 2–5.' return 'Character stills are extra views of the same person. Picture 1 is the last frame after shot 1. Permanence stills only name the lock text.'
} }
return 'Last-frame I2V does not send these stills to MiniMax as extra pictures. Continuity is the previous clip’s last frame; the names only go in the lock line.' return 'Last-frame I2V does not send these stills to MiniMax as extra pictures. Continuity is the previous clip’s last frame; the names only go in the lock line.'
}) })
@@ -2667,7 +2660,7 @@ const videoModelHint = computed(() => {
: 'PinkCherry image-to-video. Distilled LoRA stays on. Silent MP4 — LTX has no audio VAE here.' : 'PinkCherry image-to-video. Distilled LoRA stays on. Silent MP4 — LTX has no audio VAE here.'
} }
if (textToVideo.value) return 'MiniMax H3 text-to-video with native stereo audio. Shot 2+ switches to last-frame image-to-video.' if (textToVideo.value) return 'MiniMax H3 text-to-video with native stereo audio. Shot 2+ switches to last-frame image-to-video.'
if (videoWorkflow.value === 'v2') return 'MiniMax H3 image-to-video. Identity refs stay on v2 with a start still.' if (videoWorkflow.value === 'v2') return 'MiniMax H3 image-to-video. Shot 2+ starts from the last frame. Character stills are optional extras.'
return 'MiniMax H3 image-to-video with native stereo audio.' return 'MiniMax H3 image-to-video with native stereo audio.'
}) })
const frameCount = computed(() => Math.max(5, Math.floor(duration.value * fps.value))) const frameCount = computed(() => Math.max(5, Math.floor(duration.value * fps.value)))
@@ -3320,7 +3313,6 @@ onMounted(async () => {
applyEngineDefaults('ltx') applyEngineDefaults('ltx')
} }
} catch { /* ignore */ } } catch { /* ignore */ }
useIdentityRefs.value = localStorage.getItem('aigen-use-identity-refs') === 'true'
videoLoraStack.value = readStoredLoraStack('aigen-video-lora') videoLoraStack.value = readStoredLoraStack('aigen-video-lora')
imageLoraStack.value = readStoredLoraStack('aigen-image-lora') imageLoraStack.value = readStoredLoraStack('aigen-image-lora')
imageV2LoraStack.value = readStoredLoraStack('aigen-image-v2-lora') imageV2LoraStack.value = readStoredLoraStack('aigen-image-v2-lora')
@@ -3469,16 +3461,6 @@ watch(videoStart, (value) => {
} catch { /* ignore */ } } catch { /* ignore */ }
}) })
watch(useIdentityRefs, (value) => {
try {
localStorage.setItem('aigen-use-identity-refs', String(value))
} catch { /* ignore */ }
})
watch(extraPersonLock, (locked) => {
if (locked) useIdentityRefs.value = false
})
watch([globalLocks, familyPermanenceRefs, shotPermanenceRefs], () => { watch([globalLocks, familyPermanenceRefs, shotPermanenceRefs], () => {
try { try {
localStorage.setItem(LOCKS_STORE, globalLocks.value) localStorage.setItem(LOCKS_STORE, globalLocks.value)
@@ -4493,7 +4475,6 @@ async function restoreStudioJob(id: string) {
} }
if (typeof payload.sound === 'boolean') withSound.value = payload.sound if (typeof payload.sound === 'boolean') withSound.value = payload.sound
shotScriptMode.value = extensions.length > 0 shotScriptMode.value = extensions.length > 0
useIdentityRefs.value = payload.useIdentityRefs === true
identityRefs.value.forEach((_, index) => clearIdentityRef(index)) identityRefs.value.forEach((_, index) => clearIdentityRef(index))
extensionQueue.value = extensions.map(item => ({ extensionQueue.value = extensions.map(item => ({
id: crypto.randomUUID(), id: crypto.randomUUID(),
@@ -4749,7 +4730,6 @@ async function rerun(item: LibraryClip, collection = false) {
applyVideoWorkflow(target.workflow) applyVideoWorkflow(target.workflow)
if (typeof target.sound === 'boolean') withSound.value = target.sound if (typeof target.sound === 'boolean') withSound.value = target.sound
shotScriptMode.value = restoreAll shotScriptMode.value = restoreAll
useIdentityRefs.value = false
identityRefs.value.forEach((_, index) => clearIdentityRef(index)) identityRefs.value.forEach((_, index) => clearIdentityRef(index))
const ordered = chronologicalParts(parts) const ordered = chronologicalParts(parts)
globalLocks.value = ordered.find(part => part.globalLocks)?.globalLocks || target.globalLocks || '' globalLocks.value = ordered.find(part => part.globalLocks)?.globalLocks || target.globalLocks || ''
+1 -1
View File
@@ -58,7 +58,7 @@
}, },
"7": { "7": {
"inputs": { "inputs": {
"lora_name": "klein_snofs_v1_4.safetensors", "lora_name": "xaigen-klein_snofs_v1_4.safetensors",
"strength_model": 0.65, "strength_model": 0.65,
"strength_clip": 0.35, "strength_clip": 0.35,
"model": ["4", 0], "model": ["4", 0],
+1 -1
View File
@@ -43,7 +43,7 @@
}, },
"7": { "7": {
"inputs": { "inputs": {
"lora_name": "klein_snofs_v1_4.safetensors", "lora_name": "xaigen-klein_snofs_v1_4.safetensors",
"strength_model": 0.65, "strength_model": 0.65,
"strength_clip": 0.35, "strength_clip": 0.35,
"model": ["4", 0], "model": ["4", 0],
+1 -1
View File
@@ -23,7 +23,7 @@
}, },
"7": { "7": {
"inputs": { "inputs": {
"lora_name": "klein_snofs_v1_4.safetensors", "lora_name": "xaigen-klein_snofs_v1_4.safetensors",
"strength_model": 0.65, "strength_model": 0.65,
"strength_clip": 0.35, "strength_clip": 0.35,
"model": ["4", 0], "model": ["4", 0],
+1 -1
View File
@@ -76,7 +76,7 @@
}, },
"7": { "7": {
"inputs": { "inputs": {
"lora_name": "klein_snofs_v1_4.safetensors", "lora_name": "xaigen-klein_snofs_v1_4.safetensors",
"strength_model": 0.65, "strength_model": 0.65,
"strength_clip": 0.3, "strength_clip": 0.3,
"model": ["4", 0], "model": ["4", 0],
+1 -1
View File
@@ -97,7 +97,7 @@
}, },
"155": { "155": {
"inputs": { "inputs": {
"prompt": "<Picture 1> is the exact facial identity, hairstyle, body proportions and clothing of the main character. Preserve this identity with high fidelity even when the subject temporarily leaves the frame and returns.\n\n[describe the action, camera and audio here]", "prompt": "<Picture 1> is the current frame: keep this pose, clothes, place, and who is in shot. Extra stills are face, hair, and body only — not a wardrobe change.\n\n[describe the action, camera and audio here]",
"width": [ "width": [
"127", "127",
0 0
+4 -6
View File
@@ -152,9 +152,9 @@ export async function queueMiniMax(
const hasImage = Boolean(params.image?.data?.length) const hasImage = Boolean(params.image?.data?.length)
const uploading = !hasImage const uploading = !hasImage
? `Queueing ${engineName} text-to-video…` ? `Queueing ${engineName} text-to-video…`
: params.useIdentityRefs : chainIndex > 0
? (chainIndex > 0 ? 'Uploading identity stills for next shot...' : 'Uploading image to ComfyUI...') ? (params.useIdentityRefs ? 'Uploading last frame and character stills...' : 'Uploading last frame to ComfyUI...')
: (chainIndex > 0 ? 'Uploading last frame to ComfyUI...' : 'Uploading image to ComfyUI...') : 'Uploading image to ComfyUI...'
const queueing = chainIndex > 0 const queueing = chainIndex > 0
? `Queueing extension on ${engineName}...` ? `Queueing extension on ${engineName}...`
: `Queueing ${engineName} job...` : `Queueing ${engineName} job...`
@@ -393,9 +393,7 @@ export async function continueQueuedExtensions(
await ensureComfyReady(ready) await ensureComfyReady(ready)
await queueMiniMax(job, { await queueMiniMax(job, {
prompt: ext.prompt, prompt: ext.prompt,
image: identity image: { filename: 'last_frame.png', data: frame, type: 'image/png' },
? params.image
: { filename: 'last_frame.png', data: frame, type: 'image/png' },
width: params.width, width: params.width,
height: params.height, height: params.height,
steps: params.steps, steps: params.steps,
+4 -4
View File
@@ -11,13 +11,13 @@ export function identityBoilerplate(extraRefs: number | number[] = 0) {
const secondary = extras.length === 0 const secondary = extras.length === 0
? '' ? ''
: extras.length === 1 : extras.length === 1
? ` <Picture ${extras[0]}> is another view of that same person for identity only, not a second character.` ? ` <Picture ${extras[0]}> is that same person for face, hair, and body only — not a second character and not a wardrobe change.`
: ` ${extras.map(n => `<Picture ${n}>`).join(' and ')} are extra views of that same person for identity only, not additional characters.` : ` ${extras.map(n => `<Picture ${n}>`).join(' and ')} are extra views of that same person for face, hair, and body only — not additional characters and not a wardrobe change.`
return `<Picture 1> is the identity lock for a single main subject: exact face, hair, body, and clothing.${secondary} Do not spawn extra people from reference stills. If this subject leaves the frame, they must return as the same person from <Picture 1>.` return `<Picture 1> is the current frame: keep this pose, clothes, place, and who is in shot. Do not reset the outfit to an earlier still.${secondary} Do not spawn extra people from reference stills. If this subject leaves the frame, they return as the person from the extra stills, wearing whatever Picture 1 already shows.`
} }
export function looksLikeIdentityPrompt(text: string) { export function looksLikeIdentityPrompt(text: string) {
return /^<Picture 1> is the (identity lock|primary identity anchor|exact facial identity)/i.test(String(text || '').trim()) return /^<Picture 1> is the (current frame|identity lock|primary identity anchor|exact facial identity)/i.test(String(text || '').trim())
} }
export function buildIdentityPrompt(actionPrompt: string, extraRefs: number | number[] = 0) { export function buildIdentityPrompt(actionPrompt: string, extraRefs: number | number[] = 0) {
+1 -1
View File
@@ -26,7 +26,7 @@ export const IMAGE_V2_STRENGTH_MIN = 0
export const IMAGE_V2_STRENGTH_MAX = 2 export const IMAGE_V2_STRENGTH_MAX = 2
export const IMAGE_V2_STRENGTH_STEP = 0.05 export const IMAGE_V2_STRENGTH_STEP = 0.05
export const IMAGE_V2_SNOFS_LORA = 'klein_snofs_v1_4.safetensors' export const IMAGE_V2_SNOFS_LORA = 'xaigen-klein_snofs_v1_4.safetensors'
export const IMAGE_V2_CONSISTENCY_LORA = 'Flux2-Klein-9B-consistency-V2.safetensors' export const IMAGE_V2_CONSISTENCY_LORA = 'Flux2-Klein-9B-consistency-V2.safetensors'
export const IMAGE_V2_DENOISE_DEFAULT = 0.35 export const IMAGE_V2_DENOISE_DEFAULT = 0.35
export const IMAGE_V2_DENOISE_MIN = 0.15 export const IMAGE_V2_DENOISE_MIN = 0.15