Files
aigen/docs/local-video-upscale.md
T

65 lines
7.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Local video upscale
Completed videos have an **Upscale** button in Studio output and Library clip details. Its dialog defaults to **2x**, **Preserve aspect**, **Keep original FPS**, and **Faithful**. Optional values are 4x, Fit 1080p / Fit 4K, 48 / 50 / 60 FPS, and Slight sharpen. Output is H.264 MP4.
Fit means a **1920 / 3840 pixel long edge**, even for portrait or square video. The short edge follows the original ratio, rounded to the nearest even pixel for H.264. There is no crop or padding. Fit overrides the multiplicative resolution, while the selected scale remains in the requested filename. The preview uses catalog dimensions; the worker probes the actual file before processing and reports its final dimensions.
## Local installation and weights
The Windows connector starts a separate Node worker with portable **Real-ESRGAN NCNN Vulkan** and **RIFE NCNN Vulkan** binaries. Vulkan runs on the local NVIDIA GPU; there is no CUDA/Python environment dependency and no Comfy graph, Comfy start/wake call, cloud engine, or paid endpoint. The existing connector process on port 8199 must be running; the ComfyUI service does not need to run.
The first upscale automatically runs `scripts/setup-upscale.ps1`. It downloads official upstream archives over HTTPS and checks the executables/models. Download or installation failures are reported on the job. Installation can also be performed ahead of time, without using the GPU:
```powershell
powershell.exe -NoProfile -NonInteractive -ExecutionPolicy Bypass -File scripts/setup-upscale.ps1
```
Default root on this machine: `C:\Users\ianjm\Development\AIGen-Upscale` (a sibling of the app repository).
| Host variable | Default / purpose |
| --- | --- |
| `UPSCALE_ROOT` | `../AIGen-Upscale` relative to the repository; tools, models, job scratch directories |
| `UPSCALE_GPU_ID` | `0`; NCNN Vulkan device index, select the 5080 if multiple GPUs are installed |
| `UPSCALE_TILE` | `128`; Real-ESRGAN tile size, smaller tiles reduce spatial-upscale VRAM use |
Weights and binaries:
- `esrgan/models/realesrgan-x4plus.bin` and `.param`: unmodified general Real-ESRGAN x4 model. Native 4x inference is followed by Lanczos resize to the requested dimensions, including 2x.
- `rife/rife-v4*/flownet.bin` and `.param`: included RIFE v4 model. The newest available v4 directory in that package is selected.
- `esrgan/realesrgan-ncnn-vulkan.exe`, `rife/rife-ncnn-vulkan.exe`, `ffmpeg/ffmpeg.exe`, `ffmpeg/ffprobe.exe`.
- `jobs/<live-job-id>/`: uploaded input, request, durable status, worker log, encoded chunks and output. No weights or media are committed to Git. Retain this directory for failure diagnosis; inactive job directories can be removed after the website has saved the result.
Download sources: [Real-ESRGAN portable release](https://github.com/xinntao/Real-ESRGAN/releases/tag/v0.2.5.0), [RIFE portable release](https://github.com/nihui/rife-ncnn-vulkan/releases/tag/20221029), [Gyan FFmpeg essentials](https://www.gyan.dev/ffmpeg/builds/).
## Queue and worker call
The dialog posts `{scale, target, fps, enhance}` to `/api/library/clips/<source-id>/upscale`. The authenticated route resolves the existing asset ID to its owned library MP4 on disk, verifies folder access, rejects duplicate work for that clip, and enqueues one Studio video row with `payload.upscale`. Arbitrary input URLs and filesystem paths are not accepted from browsers.
The Studio dispatcher calls `startUpscaleJob(item)` in `server/utils/videoUpscale.ts` under the existing shared GPU reservation. It streams the source file to `PUT /upscale/jobs/<live-id>/input`, then posts `{id, scale, target, fps, enhance}` to `/upscale/jobs` on the local connector. The host executes:
```text
node scripts/upscale-worker.mjs <UPSCALE_ROOT>/jobs/<live-id>/request.json
```
These requests reuse `COMFY_CONTROL_URL` and `COMFY_CONTROL_TOKEN` (or their existing Nuxt runtime overrides) **only as the connector address and authentication**. No Comfy service is launched. GPU reservation prevents overlap with generation/training jobs. New generations can still be queued while an upscale is running. Upscales have their own clip status and do not take over the generation input form.
The worker decodes approximately two seconds plus one overlapping source frame per chunk, spatially upscales with tiled Real-ESRGAN, resizes to target, optionally slightly sharpens, then runs RIFE if a numeric output FPS was selected. RIFE uses UHD mode to reduce memory pressure. Integer oversampling followed by FFmpeg frame selection supports 48, 50 and 60 FPS, including fractional source frame rates. Chunk frame counts use global timeline boundaries to avoid cumulative audio drift. Chunks are encoded once, concatenated without re-encoding, and returned as **one job and one file** for short clips or 60–120 second masters.
Exact final remux (the worker supplies the actual source duration and absolute paths):
```text
ffmpeg -v error -f concat -safe 1 -i concat.txt -i input.mp4 -map 0:v:0 -map 1:a? -c:v copy -c:a copy -map_metadata 1 -t <source-duration-seconds> -movflags +faststart output.mp4
```
`1:a?` makes audio optional; `-c:a copy` preserves original encoded audio. There is no interpolation or time stretching of audio. An unsupported source/audio stream fails cleanly rather than silently discarding audio. Source resolution, target resolution, duration, engine and elapsed milliseconds appear in the worker log/status and website server log.
The website writes a collision-safe sibling `{stem}_up2x.mp4` or `{stem}_up4x.mp4`, then `-2`, `-3`, etc. Canonical library sources are named `video.mp4`, so their sibling starts with `video_up2x.mp4`. A separate catalog entry in the same folder stores `upscaledFromClipId` and `upscaleJobId`, with an “Open original” link. This deliberately uses a distinct derivative relationship instead of the extension-specific `parentClipId` / family chain, so the existing last-frame extension and stitching logic continues to choose its original clips. Large outputs are copied from disk rather than loaded into one JavaScript buffer.
Jobs survive website restarts through `<LIBRARY_DIR>/upscale-jobs/*.json`. A lost HTTP response is checked by job ID instead of resubmitting inference. Host restart or cancellation terminates the worker tree; originals are retained. OOM, missing weights and bad files become failed jobs. No fallback generation, engine, or cloud request exists.
## Rollout and validation
Deploy the site and restart the Windows connector when the shared GPU is idle. The connector must have the adjacent `upscale-host.mjs`, `upscale-worker.mjs`, `setup-upscale.ps1`, and `../shared/video-upscale.mjs` files from the same checkout. A Coolify deployment alone does not reload the Windows connector.
`npm test` covers defaults, aspect ratios, cancel-before-start, duplicate submissions, OOM release, queue ownership and a mocked two-minute fractional-FPS pipeline. With `UPSCALE_TEST_FFMPEG` pointing to an FFmpeg executable (or `.data/video-test-tools/ffmpeg.exe` present), a CPU-only synthetic test substitutes both neural engines and checks the real chunk encoding/remux, output frame count and bit-identical copied audio. It never runs Real-ESRGAN/RIFE inference. Actual 5080 inference performance and visual quality require a later user-run upscale.