Exclusive llama-server load/unload on the 5080, queued Describe UI with resultText, Copy, and Use as prompt. Co-authored-by: Cursor <cursoragent@cursor.com>
1.6 KiB
Studio 2 image→text caption (Qwen2.5-VL NSFW Caption V4 GGUF)
Caption jobs run on the Windows GPU host via llama-server (llama.cpp vision / mmproj). Not Comfy, not Klein, not JoyCaption.
Weights (~7GB — two files only)
powershell -ExecutionPolicy Bypass -File scripts\setup-caption.ps1
Installs into the Shared models tree:
%LOCALAPPDATA%\Comfy-Desktop\ComfyUI-Shared\models\caption\qwen25vl-7b-nsfw-v4\
| File | Role |
|---|---|
Qwen2.5-VL-7B-NSFW-Caption-V4.Q5_K_M.gguf |
Language model |
Qwen2.5-VL-7B-NSFW-Caption-V4.mmproj-f16.gguf |
Vision projector |
Do not download the full 16.6GB safetensors repo.
Runtime needs ~8–10GB VRAM. Exclusive GPU: host stops Comfy (and waits if YuE2/upscale is busy), loads the GGUF for one shot, then kills llama-server (keep_alive 0).
Host agent
Requires llama-server on PATH (winget install ggml.llamacpp) or CAPTION_LLAMA_SERVER.
Restart the Comfy host agent after setup. Confirm:
GET http://127.0.0.1:8199/caption/status → { configured, busy, backend: "llama.cpp" }
POST /caption { imagePath, style } → load → { text } → unload
POST /caption/jobs + PUT …/input → async job used by Studio 2
Styles: descriptive | klein_prompt | delta | tags.
App
POST /api/studio-2/caption with { folderId, stillId|sourcePath, captionStyle } queues a Studio job. Bench Describe button and the caption task pill both enqueue. Result stores resultText on the studio2 job and optionally writes {still}.txt beside the library file.