Store the Edit TextEncode prompt on library stills, not a Start still caption.

Edit + Enhance always leads with the keep/<image1> stanza and typed change; PE describe dumps are dropped, and Typed stays promptRaw.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
Towsty
2026-09-26 15:14:20 -05:00
co-authored by Cursor
parent c46e827c68
commit 58c5458c87
9 changed files with 177 additions and 46 deletions
@@ -58,6 +58,10 @@ Before emitting JSON, reason through: what the image(s) actually contain (includ
For every input, the rewritten instruction MUST use <image1>, <image2>, … to refer to each input image. Single-image edits still use <image1>. Never write "the photo", "the image", "the woman in the picture" as a substitute for the tag. The first Start still is always <image1>.
The rewritten prompt is an EDIT INSTRUCTION, not a description of the input image.
Lead with the operation. Mention <image1> in the first sentence.
Do not write "The image is a photograph of…". That format is for text-to-image only.
State each image's role explicitly — which one is the canvas whose composition and untargeted content survive, and which supply material to transfer — and say what is taken from each. For scene generation with no canvas (合影/合照 and the like), all images serve as identity sources. Describe every referenced image individually; never compress several into a range or a group to avoid describing them one by one.
## Output Size Determination