# Studio 2 image→text caption (Qwen2.5-VL NSFW Caption V4 GGUF) Caption jobs run on the Windows GPU host via **llama-server** (llama.cpp vision / mmproj). Not Comfy, not Klein, not JoyCaption. ## Weights (~7GB — two files only) ```powershell powershell -ExecutionPolicy Bypass -File scripts\setup-caption.ps1 ``` Installs into the Shared models tree: `%LOCALAPPDATA%\Comfy-Desktop\ComfyUI-Shared\models\caption\qwen25vl-7b-nsfw-v4\` | File | Role | | --- | --- | | `Qwen2.5-VL-7B-NSFW-Caption-V4.Q5_K_M.gguf` | Language model | | `Qwen2.5-VL-7B-NSFW-Caption-V4.mmproj-f16.gguf` | Vision projector | Do **not** download the full 16.6GB safetensors repo. Runtime needs ~8–10GB VRAM. Exclusive GPU: host stops Comfy (and waits if YuE2/upscale is busy), loads the GGUF for one shot, then kills llama-server (`keep_alive 0`). ## Host agent Requires `llama-server` on PATH (`winget install ggml.llamacpp`) or `CAPTION_LLAMA_SERVER`. Restart the Comfy host agent after setup. Confirm: ```text GET http://127.0.0.1:8199/caption/status → { configured, busy, backend: "llama.cpp" } POST /caption { imagePath, style } → load → { text } → unload POST /caption/jobs + PUT …/input → async job used by Studio 2 ``` Styles: `descriptive` | `klein_prompt` | `delta` | `tags`. ## App `POST /api/studio-2/caption` with `{ folderId, stillId|sourcePath, captionStyle }` queues a Studio job. Bench **Describe** button and the **caption** task pill both enqueue. Result stores `resultText` on the studio2 job and optionally writes `{still}.txt` beside the library file.