Files
aigen/docs/caption.md
T
TowstyandCursor 1b311d6529 Add Studio 2 caption jobs via Qwen2.5-VL GGUF on the host agent.
Exclusive llama-server load/unload on the 5080, queued Describe UI with resultText, Copy, and Use as prompt.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-30 22:06:10 -05:00

1.6 KiB
Raw Blame History

Studio 2 image→text caption (Qwen2.5-VL NSFW Caption V4 GGUF)

Caption jobs run on the Windows GPU host via llama-server (llama.cpp vision / mmproj). Not Comfy, not Klein, not JoyCaption.

Weights (~7GB — two files only)

powershell -ExecutionPolicy Bypass -File scripts\setup-caption.ps1

Installs into the Shared models tree:

%LOCALAPPDATA%\Comfy-Desktop\ComfyUI-Shared\models\caption\qwen25vl-7b-nsfw-v4\

File Role
Qwen2.5-VL-7B-NSFW-Caption-V4.Q5_K_M.gguf Language model
Qwen2.5-VL-7B-NSFW-Caption-V4.mmproj-f16.gguf Vision projector

Do not download the full 16.6GB safetensors repo.

Runtime needs ~8–10GB VRAM. Exclusive GPU: host stops Comfy (and waits if YuE2/upscale is busy), loads the GGUF for one shot, then kills llama-server (keep_alive 0).

Host agent

Requires llama-server on PATH (winget install ggml.llamacpp) or CAPTION_LLAMA_SERVER.

Restart the Comfy host agent after setup. Confirm:

GET http://127.0.0.1:8199/caption/status  →  { configured, busy, backend: "llama.cpp" }
POST /caption  { imagePath, style }       →  load → { text } → unload
POST /caption/jobs + PUT …/input          →  async job used by Studio 2

Styles: descriptive | klein_prompt | delta | tags.

App

POST /api/studio-2/caption with { folderId, stillId|sourcePath, captionStyle } queues a Studio job. Bench Describe button and the caption task pill both enqueue. Result stores resultText on the studio2 job and optionally writes {still}.txt beside the library file.