Exclusive llama-server load/unload on the 5080, queued Describe UI with resultText, Copy, and Use as prompt. Co-authored-by: Cursor <cursoragent@cursor.com>
41 lines
1.6 KiB
Markdown
41 lines
1.6 KiB
Markdown
# Studio 2 image→text caption (Qwen2.5-VL NSFW Caption V4 GGUF)
|
||
|
||
Caption jobs run on the Windows GPU host via **llama-server** (llama.cpp vision / mmproj). Not Comfy, not Klein, not JoyCaption.
|
||
|
||
## Weights (~7GB — two files only)
|
||
|
||
```powershell
|
||
powershell -ExecutionPolicy Bypass -File scripts\setup-caption.ps1
|
||
```
|
||
|
||
Installs into the Shared models tree:
|
||
|
||
`%LOCALAPPDATA%\Comfy-Desktop\ComfyUI-Shared\models\caption\qwen25vl-7b-nsfw-v4\`
|
||
|
||
| File | Role |
|
||
| --- | --- |
|
||
| `Qwen2.5-VL-7B-NSFW-Caption-V4.Q5_K_M.gguf` | Language model |
|
||
| `Qwen2.5-VL-7B-NSFW-Caption-V4.mmproj-f16.gguf` | Vision projector |
|
||
|
||
Do **not** download the full 16.6GB safetensors repo.
|
||
|
||
Runtime needs ~8–10GB VRAM. Exclusive GPU: host stops Comfy (and waits if YuE2/upscale is busy), loads the GGUF for one shot, then kills llama-server (`keep_alive 0`).
|
||
|
||
## Host agent
|
||
|
||
Requires `llama-server` on PATH (`winget install ggml.llamacpp`) or `CAPTION_LLAMA_SERVER`.
|
||
|
||
Restart the Comfy host agent after setup. Confirm:
|
||
|
||
```text
|
||
GET http://127.0.0.1:8199/caption/status → { configured, busy, backend: "llama.cpp" }
|
||
POST /caption { imagePath, style } → load → { text } → unload
|
||
POST /caption/jobs + PUT …/input → async job used by Studio 2
|
||
```
|
||
|
||
Styles: `descriptive` | `klein_prompt` | `delta` | `tags`.
|
||
|
||
## App
|
||
|
||
`POST /api/studio-2/caption` with `{ folderId, stillId|sourcePath, captionStyle }` queues a Studio job. Bench **Describe** button and the **caption** task pill both enqueue. Result stores `resultText` on the studio2 job and optionally writes `{still}.txt` beside the library file.
|