Files
aigen/docs/yuegp-replacement.md
T

104 lines
7.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Standalone YuEGP replacement
The app's `yue` music engine now means **standalone deepbeepmeep/YuEGP**. It never submits a Comfy graph. ACE-Step and ACE-Step 1.5 retain their existing Comfy workflows. No image/video model or extension workflow was changed.
## Defaults and controls
- Profile **1**: unquantized YuEGP path. Upstream loads Stage 1 in BF16 and Stage 2 in FP16; the UI does not mislabel these as a single FP16 mode.
- Profile **3** is a manual setting only. No automatic OOM fallback, retry into a different profile, or profile 2 path.
- **60 seconds**, **one non-empty lyric section**, **6,000 maximum new tokens**. Duration maps to 100 tokens/second; actual audio duration depends on the model's end token. Supported single-section range is 30–150 seconds (below the 16,384-token context limit).
- English CoT Stage 1: `m-a-p/YuE-s1-7B-anneal-en-cot`.
- Stage 2: `m-a-p/YuE-s2-1B-general`.
- ICL and dual-track prompting off. Instrumental-only generation remains available through ACE.
- Compile **off**. The app never passes `--compile`. Explicit command-line opt-in checks `import triton` before enabling it.
- Lyrics, genre tags, seed (including zero), destination folder/name, duration, and profile are carried through the music job. More than one non-empty lyric section is rejected rather than silently dropping lyrics. Existing numbered/hyphenated headings are normalized for upstream tokenization.
## Exact call path
The UI sends `POST /api/generate/music` with:
```json
{"engine":"yue","yueProfile":1,"duration":60,"seed":42,"tags":"pop, warm vocals","lyrics":"[Verse 1]\nYour lyrics here","folderId":"<library-folder>","name":"My song"}
```
The studio queue calls `startMusicJob(params)`, which routes `engine === 'yue'` directly to `startYueGpJob(params)`. Under the existing shared GPU reservation, that sends:
```text
POST <COMFY_CONTROL_URL>/yuegp/jobs
{ id, profile: 1, duration: 60, seed, tags, lyrics }
```
The device-local host launches this argument array (no shell, hidden Windows process):
```text
<YUEGP_ROOT>/.venv/Scripts/python.exe -u <aigen>/scripts/yuegp-worker.py --root <YUEGP_ROOT> --request <job-dir>/request.json --profile 1
```
The worker calls the pinned YuEGP functions:
```python
offload.profile({'transformer': model, 'stage2': model2},
profile_no=1, quantizeTransformer=False, compile=False, verboseLevel=1)
stage1_inference(tags, lyrics, 1, 6000, seed, state, callback)
stage2_inference(model2, stems, output_dir, batch_size=20, state=state, callback=callback)
```
The inference functions are loaded from upstream `inference/gradio_server.py` using an explicit AST function list and source checksum. No Gradio import, UI, server, CLI initialization, or upstream default profile executes. The upstream standalone `infer.py` is not used: its missing `sdpa` argument and single-section loop differ from the maintained functions.
## Host setup
Run `scripts/setup-yuegp.ps1 -Uv <path-to-uv.exe>` on the GPU computer. The default sibling checkout is `C:\Users\ianjm\Development\YuEGP`, pinned to revision `2d72ff734b7a127324353c0dcd0f95ca4cc0b797`. Setup creates its own Python 3.10 environment with PyTorch 2.7.1/CUDA 12.8 for the RTX 5080. It installs YuEGP's own Transformers optimization files only into that environment. It does not modify Comfy's Python, install Gradio, or commit weights.
Existing Stage 1/2 weights and codec weights are reused from the local model directory recorded in `aigen-ready.json`. Missing model weights use their HF IDs. Optional host environment variables:
| Variable | Purpose |
| --- | --- |
| `YUEGP_ROOT` | Standalone YuEGP checkout |
| `YUEGP_PYTHON` | Isolated worker Python executable |
| `YUEGP_JOBS_DIR` | Durable per-job logs, requests, intermediates, and WAV output |
| `YUEGP_MODELS_DIR` | Parent directory containing Stage 1/2 model folders |
| `YUEGP_STAGE1_MODEL` / `YUEGP_STAGE2_MODEL` | Explicit local model paths or HF IDs |
The deployed Nuxt app uses its existing `COMFY_CONTROL_URL` and token. Restart the host connector after installing the new scripts. This does not require changing any image/video workflow.
## Queue, cancellation, and output
The host verifies Comfy's queue is empty, stops the idle Comfy process to release its cached VRAM, and launches YuEGP under the same shared GPU reservation. Comfy mutations and wake requests are blocked while YuEGP owns the GPU. Cancellation waits for worker exit. Expired reservations stop the worker; a parent-process watchdog prevents orphan GPU jobs after a host crash.
Status is read from `/yuegp/jobs/<id>` and audio from `/yuegp/jobs/<id>/audio`. Token counts and stage names are real YuEGP callbacks; the percentage describes the current stage, not an ETA. The app saves the final WAV into its normal library and releases the studio queue slot. Duplicate submissions use the same ID. Durable app records restore pending jobs before queue repair on restart, and saved audio is deduplicated by job ID. Errors preserve intermediates and never fall back to Comfy.
## Removed
- `server/assets/workflow_yue.json`
- `scripts/patch_yue_cancellation.py`
- `tests/test_yue_cancellation.py`
- Local untracked `scripts/patch_comfyui_yue_windows.ps1`
- YuE graph builder/node checks, hardcoded MMGP profile 2 payloads, obsolete model environment settings, and the Comfy-node model-copy installer branch.
Historical library metadata may still recognize old YuE outputs. That is read-only compatibility, not a generation backend. Old host custom-node files are not used by this integration.
## Added
- `scripts/yuegp-worker.py`
- `scripts/yuegp-host.mjs`
- `scripts/setup-yuegp.ps1`
- `server/utils/yueGp.ts`
- `server/plugins/00-resume-yuegp.ts`
- `tests/yuegp.test.mjs`
- `tests/test_yuegp_worker.py`
- This document.
UI, queue, host reservation/proxy, library metadata, and music API files were updated to connect these pieces. Existing music engine key `yue` is preserved for compatibility.
Upstream reference: [deepbeepmeep/YuEGP](https://github.com/deepbeepmeep/YuEGP/tree/2d72ff734b7a127324353c0dcd0f95ca4cc0b797). Its advertised timings are not a benchmark of this 5080; local validation is recorded below.
## Validation
- 45 JavaScript tests passed, including shared GPU coordination, YuEGP process lifecycle, no fallback, progress parsing, and unchanged ACE graph construction.
- Two CPU-only Python regression tests passed for lyric preservation and explicit CPU decoder construction/loading. These do not import PyTorch or run inference.
- Production build passed.
- Source scan finds no live Comfy YuE generation graph. `server/utils/comfy.ts` still recognizes `YUE_Stage_A_Sampler` only when reading historical metadata.
- The earlier test failed during decoder construction: MMGP sets the default PyTorch device to CUDA. The adapter now defers decoder creation until the transformer stages finish, releases their VRAM, and constructs the decoder in an explicit CPU context. It also writes WAV directly instead of depending on a Windows MP3 backend.
- No generation test was run after that fix, at the user's request. Successful song generation, peak VRAM, and wall time remain unverified.
- This isolated environment currently has no FlashAttention or Triton. It uses upstream SDPA with compilation off. The optimized FlashAttention cache path in YuEGP is therefore not active; profile 1 remains selected and unquantized, but advertised four-minute timings must not be assumed for this installation.