Add Studio 2 caption jobs via Qwen2.5-VL GGUF on the host agent.

Exclusive llama-server load/unload on the 5080, queued Describe UI with resultText, Copy, and Use as prompt.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
Towsty
2026-09-30 22:06:10 -05:00
co-authored by Cursor
parent 1f10082087
commit 1b311d6529
14 changed files with 1285 additions and 41 deletions
+40
View File
@@ -0,0 +1,40 @@
# Studio 2 image→text caption (Qwen2.5-VL NSFW Caption V4 GGUF)
Caption jobs run on the Windows GPU host via **llama-server** (llama.cpp vision / mmproj). Not Comfy, not Klein, not JoyCaption.
## Weights (~7GB — two files only)
```powershell
powershell -ExecutionPolicy Bypass -File scripts\setup-caption.ps1
```
Installs into the Shared models tree:
`%LOCALAPPDATA%\Comfy-Desktop\ComfyUI-Shared\models\caption\qwen25vl-7b-nsfw-v4\`
| File | Role |
| --- | --- |
| `Qwen2.5-VL-7B-NSFW-Caption-V4.Q5_K_M.gguf` | Language model |
| `Qwen2.5-VL-7B-NSFW-Caption-V4.mmproj-f16.gguf` | Vision projector |
Do **not** download the full 16.6GB safetensors repo.
Runtime needs ~8–10GB VRAM. Exclusive GPU: host stops Comfy (and waits if YuE2/upscale is busy), loads the GGUF for one shot, then kills llama-server (`keep_alive 0`).
## Host agent
Requires `llama-server` on PATH (`winget install ggml.llamacpp`) or `CAPTION_LLAMA_SERVER`.
Restart the Comfy host agent after setup. Confirm:
```text
GET http://127.0.0.1:8199/caption/status → { configured, busy, backend: "llama.cpp" }
POST /caption { imagePath, style } → load → { text } → unload
POST /caption/jobs + PUT …/input → async job used by Studio 2
```
Styles: `descriptive` | `klein_prompt` | `delta` | `tags`.
## App
`POST /api/studio-2/caption` with `{ folderId, stillId|sourcePath, captionStyle }` queues a Studio job. Bench **Describe** button and the **caption** task pill both enqueue. Result stores `resultText` on the studio2 job and optionally writes `{still}.txt` beside the library file.