Avoid YuEGP pinned-memory exhaustion on 32GB Windows hosts

This commit is contained in:
Towsty
2026-09-08 07:04:13 -05:00
parent 7665285a0d
commit 4c97980646
3 changed files with 35 additions and 2 deletions
+4
View File
@@ -101,3 +101,7 @@ Upstream reference: [deepbeepmeep/YuEGP](https://github.com/deepbeepmeep/YuEGP/t
- The earlier test failed during decoder construction: MMGP sets the default PyTorch device to CUDA. The adapter now defers decoder creation until the transformer stages finish, releases their VRAM, and constructs the decoder in an explicit CPU context. It also writes WAV directly instead of depending on a Windows MP3 backend.
- No generation test was run after that fix, at the user's request. Successful song generation, peak VRAM, and wall time remain unverified.
- This isolated environment currently has no FlashAttention or Triton. It uses upstream SDPA with compilation off. The optimized FlashAttention cache path in YuEGP is therefore not active; profile 1 remains selected and unquantized, but advertised four-minute timings must not be assumed for this installation.
## Windows startup memory correction
The later profile-1 run reached Stage 1 but failed before its first forward pass. MMGP logged a failed pinned-host-memory allocation after reserving about 13GB, then CUDA reported OOM in special-token preparation. On Windows hosts with less than 48GB system RAM, the adapter now passes `pinnedMemory=False` to `offload.profile`. Profile 1, unquantized weights, default VRAM budgets, 60s, and compile-off remain unchanged. This disables transfer-staging RAM pinning, not profile 1, and does not switch to profile 2 or 3. CUDA synchronization and free/allocated GPU and available RAM diagnostics now bracket initialization and Stage 1 completion. No generation was run to validate this correction; subsequent inference may still encounter a separate VRAM limit.