MS-SSIM¶
MS-SSIM (Multi-Scale Structural Similarity Index Measure) extends SSIM to a multi-resolution pyramid, providing a perceptual similarity metric that is robust to viewing distance and display resolution variation. Each scale captures structural information at a different spatial frequency. The score is in [0, 1], higher is better, and 1.0 means identical frames.
The fork ships one CPU extractor, float_ms_ssim. It uses the IQA library's Gaussian-window floating-point implementation with a 5-scale Laplacian pyramid (Wang et al. 2004), and it is the extractor invoked when VMAF model JSON files reference "float_ms_ssim". The CLI also accepts integer_ms_ssim as an alias for it.
How to run¶
# Luma-only MS-SSIM (default)
build/tools/vmaf \
--reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--no_prediction --feature float_ms_ssim --output /dev/stdout
# Per-channel MS-SSIM (luma + Cb + Cr)
build/tools/vmaf \
--reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--no_prediction --feature float_ms_ssim=enable_chroma=true --output /dev/stdout
The minimum supported input resolution is 176x176 (see Minimum resolution).
Output features¶
| Feature name | Description | Condition |
|---|---|---|
float_ms_ssim | MS-SSIM on the luma (Y) plane | Always |
float_ms_ssim_cb | MS-SSIM on the Cb (U) chroma plane | enable_chroma=true only |
float_ms_ssim_cr | MS-SSIM on the Cr (V) chroma plane | enable_chroma=true only |
float_ms_ssim_l_scale0-4 | Per-scale luminance component | enable_lcs=true, luma only |
float_ms_ssim_c_scale0-4 | Per-scale contrast component | enable_lcs=true, luma only |
float_ms_ssim_s_scale0-4 | Per-scale structure component | enable_lcs=true, luma only |
Options¶
| Option | Type | Default | Effect |
|---|---|---|---|
enable_chroma | bool | false | Emit per-plane _cb and _cr scores in addition to luma. YUV400P sources are always luma-only. |
enable_lcs | bool | false | Emit per-scale luminance, contrast, and structure intermediate components for the luma plane. |
enable_db | bool | false | Report the luma MS-SSIM score as dB (\(-10 \log_{10}(1 - \mathrm{score})\)). |
clip_db | bool | false | Cap the dB score at a ceiling derived from the frame geometry. Only meaningful when enable_db=true. |
clip_db sets \(\mathrm{max\_db} = \left\lceil 10 \log_{10}(\mathrm{peak}^2 / \mathrm{mse}) \right\rceil\) with \(\mathrm{mse} = 0.5 / (w\,h)\). It is a ceiling on the dB output, not a clamp on the linear score, and it also defines what a perfect match reports: score >= 1.0 returns max_db rather than +Inf.
GPU clip_db scores before ADR-1221 were wrong
Up to and including v3.2.1 the CUDA, SYCL and HIP twins read clip_db as a clamp on the linear score (\([0, 1]\), then \(-10 \log_{10}(1 - \mathrm{score})\) with no ceiling) and carried no max_db at all. Scoring an identical reference/distorted pair returned +Inf, and every high-similarity pair returned an uncapped dB value, so clip_db did not clip. Fixed per ADR-1221. Re-measure any GPU MS-SSIM dB score taken with clip_db set. The Metal twin has been brought to full parity by ADR-1334, implementing enable_db, clip_db with max_db ceiling, and enable_chroma.
Minimum resolution¶
The 5-level 11-tap pyramid needs every scored plane to be at least 176x176. Smaller inputs fall below the 11-tap Gaussian kernel footprint and are rejected with an error at init time (Netflix#1414 / ADR-0153).
With enable_chroma the rule includes the subsampled planes:
- 4:2:0. Plane allocation uses ceil subsampling, so the exact luma minimum is 351x351 (352x352 is the next even-sized input), not 176x176.
- 4:2:2. The luma minimum is 351x176.
- YUV400P. There is no chroma: the option is ignored and luma is scored.
- Too small. The CPU extractor and every twin refuse an input whose scored plane is smaller, at init, and name the chroma size they measured. A run on this repository's 576x324 4:2:0 Netflix fixture (288x162 chroma) with
enable_chromatherefore writes no scores on any backend.
Before ADR-1299 the CPU extractor checked luma only, and a 4:2:0 input in between died mid-run with error: scale below 1x1! on stdout and no output file.
GPU twins¶
Every GPU twin takes the CPU extractor's four options and computes what they ask for on the device (read from each twin's own options[] table):
| Twin | Backend | enable_lcs | enable_db | clip_db | enable_chroma |
|---|---|---|---|---|---|
float_ms_ssim_cuda | CUDA | yes | yes | yes | yes, computed on the GPU (3 planes), since 2026-10-03 |
float_ms_ssim_sycl | SYCL | yes | yes | yes | yes, computed on the GPU (3 planes) |
integer_ms_ssim_hip | HIP | yes | yes | yes | yes, computed on the GPU (3 planes), since 2026-10-03 |
float_ms_ssim_metal | Metal | yes | yes | yes | yes, computed on the GPU (3 planes) |
With enable_chroma=true each twin gives every scored plane its own geometry, pyramid and score storage, runs the luma pipeline on it and writes float_ms_ssim_cb and float_ms_ssim_cr. Every twin lists both among the features it provides, so a model that asks for them is served on the GPU. The SYCL twin got the option with ADR-1299 and the Metal twin with ADR-1334.
HIP runs with enable_chroma before 2026-10-03 have no chroma scores
integer_ms_ssim_hip accepted the option, scored luma only and wrote neither float_ms_ssim_cb nor float_ms_ssim_cr, without a warning. The CUDA twin had no such option, so the same request ran the CPU extractor and its scores were right. Re-run any HIP measurement that needs the chroma scores.
Selecting the twin from the command line¶
--backend cuda, --backend sycl or --backend hip with --feature float_ms_ssim=enable_chroma=true runs that backend's twin (the JSON output's feature_backends names it); --backend cpu runs the CPU extractor. A model that names the features reaches the twin the same way. The Vulkan backend was removed in ADR-0726.
vmaf --reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--backend cuda --feature float_ms_ssim=enable_chroma=true \
--no_prediction --json --output ms_ssim.json --precision max
Tracked as T-MS-SSIM-GPU-CHROMA-OPTION-DRIFT-2026-09-06 in docs/state.md.
Agreement with the CPU¶
| Twin | vs CPU | ADR | Evidence (fragments in scripts/ci/exact_twins.d/) |
|---|---|---|---|
float_ms_ssim_cuda | Exact | ADR-1403, ADR-1465 | float_ms_ssim.cuda, float_ms_ssim_lcs.cuda, float_ms_ssim_chroma.cuda |
float_ms_ssim_sycl | Exact | ADR-1414, ADR-1466 | float_ms_ssim.sycl, float_ms_ssim_lcs.sycl, float_ms_ssim_chroma.sycl |
integer_ms_ssim_hip | Exact | ADR-1403, ADR-1437, ADR-1438 | float_ms_ssim.hip, float_ms_ssim_lcs.hip, float_ms_ssim_chroma.hip |
float_ms_ssim_metal | Tolerance 5e-5 (FEATURE_TOLERANCE); still uses the old arithmetic | none | none |
All three exact twins reproduce the CPU's arithmetic and add every per-scale sum in the CPU's order. The cost of that, in milliseconds per frame (CPU extractor on the same host for reference):
| Twin | Resolution | Time per frame | Before |
|---|---|---|---|
| SYCL (Arc A380, ADR-1414) | 3840x2160 | 42.6 ms | 31.4 ms |
| SYCL (Arc A380, ADR-1414) | 576x324 | 0.84 to 1.20 ms | not recorded |
| SYCL (Arc A380, ADR-1466) | 3840x2160 | 75.9 ms | 44.7 ms |
| SYCL (Arc A380, ADR-1466) | 1920x1080 | 18.2 ms | 11.5 ms |
| SYCL (Arc A380, ADR-1466) | 576x324 | 1.92 ms | 1.16 ms |
| CPU extractor | 3840x2160, same host | 117 ms |
The ADR-1466 SYCL twin also needs 219 MB of device and pinned host memory at 3840x2160 (SYCL backend).
How the CUDA and SYCL twins compute it¶
- The decimate fuses each tap as
ms_ssim_decimate.cdoes. - The Gaussian window sums are fp32 products added as the CPU's
doublesum is. - The luminance, contrast and structure terms use the CPU's operand types.
- The CUDA twin does this in
double. The SYCL twin, which may not usedoubleon the device, carries those values as exact pairs of floats and adds the frame sums in 64-bit fixed point.
Measured agreement of float_ms_ssim_sycl (ADR-1414)¶
Arc A380, --precision max against --backend cpu, frames identical and the largest difference:
| Fixture | Before | After |
|---|---|---|
| Netflix 576x324, 48 frames | 0 of 48, 6.9e-8 | 48 of 48 |
| Checkerboard 1920x1080, 1 px shift, 3 frames | 0 of 3, 1.06e-6 | 3 of 3 |
| Checkerboard 1920x1080, 10 px shift, 3 frames | 0 of 3, 2.98e-6 | 3 of 3 |
| BBB 3840x2160 | 0 of 50, 1.23e-6 | 199 of 200, 1.1e-16 |
With enable_lcs all 15 per-scale means are the CPU's on every frame (50 BBB frames), and so are float_ms_ssim_cb / float_ms_ssim_cr with enable_chroma, and the score at 10, 12 and 16 bits.
The reference in this table is a GCC build. The one BBB frame that differs does so in the last bit of the double: the twin's host combine calls the pow() of its own Intel-compiler build. Against the CPU extractor of its own binary the twin is identical on every frame. With enable_db the same holds for log10() (3.6e-15 on 3 of 48 Netflix frames against a GCC build).
Exact by construction (ADR-1466)¶
The match after ADR-1414 was exact in practice, not by construction. The CPU adds each frame's terms into a running double and the twins added them in another order. The rounding of each per-scale mean to float absorbs the difference unless the mean sits next to a rounding boundary.
Such frames exist (T-GPU-FLOAT-SSIM-FRAME-SUM-ORDER-2026-10-02). On a 176x176 noise pair float_ms_ssim_sycl returned 0.9884905219078064 for float_ms_ssim_l_scale0 where the CPU returns 0.9884904623031616, one float step.
Since 2026-10-02 the SYCL twin stores every window's l, c and s of every scale, with l and c as the CPU's double values, and adds them on the host in the CPU's order. It returns the CPU's value on that pair. Against a GCC build of the CPU extractor on an Arc A380: float_ms_ssim on 138 of 138 frames and 2208 of 2208 enable_lcs values.
The chroma planes¶
With enable_chroma the twins run the same kernels and host sums on each chroma plane, so the chroma scores are exact for the same reasons the luma score is. Measured on 2026-10-03 at --precision max against the CPU extractor of a GCC build, with each option set (enable_chroma alone, with enable_lcs, with enable_db, with enable_db and clip_db):
| Fixture | Frames | CUDA (RTX 4090) | HIP (gfx1036) |
|---|---|---|---|
| Checkerboard 1920x1080 4:2:0, 1 px and 10 px shift | 3 + 3 | identical | identical |
| Netflix 576x324 4:2:2 10-bit | 48 | identical | identical |
| Netflix 576x324 4:4:4, 8 and 10 bit | 48 + 48 | identical | identical |
| Netflix 576x324 4:4:4, reference against itself | 48 | identical | identical |
| BBB 1920x1080 4:4:4 | 24 | identical | identical |
| BBB 3840x2160 4:2:0 | 30 | identical | identical |
| Netflix 576x324 4:2:0 (288x162 chroma) | 48 | refused, as the CPU | refused, as the CPU |
That is 6 804 of 6 804 values on each of CUDA and HIP, the feature_backends of every run naming the twin.
- SYCL. The twin (Arc A380, build of
master9aa990455) returns the CPU extractor of its own binary on the same runs, 5 508 of 5 508 values. Against the GCC build 52 of 6 804 values differ by at most 2.1e-14, all from the Intel math library'spow()andlog10()in its host combine. - Parity gate. The
float_ms_ssim_chromacell compares the three scores with tolerance 0 on CUDA, SYCL and HIP. On a fixture whose chroma is below 176 pixels it is reportedSKIPwith the reason.
The chroma planes cost what their area costs: a 4:2:0 frame scores 1.5 times the luma area, a 4:4:4 frame 3 times. Milliseconds per frame, (t(N) - t(2)) / (N - 2) through the vmaf tool, median of 3:
| Input | CUDA luma | CUDA with chroma | HIP luma | HIP with chroma |
|---|---|---|---|---|
| BBB 3840x2160 4:2:0, N = 42 | 33.0 | 46.5 | 186.7 | 257.8 |
| BBB 1920x1080 4:4:4, N = 24 | 9.7 | 22.6 | 39.1 | 119.2 |
Device and pinned host memory grow by the same factor: every plane keeps its own pyramid and per-window terms.
History¶
- 2026-10-06. The CUDA twin builds level 0 of its pyramids on the device (
picture_copy()'s arithmetic, scores unchanged): no picture plane goes to the host and back, and the per-plane pinned staging buffers are gone. The CUDA timings above were measured before this change. - 2026-10-03. CUDA and HIP twins compute
enable_chromaon the device (HIP before this date scored luma only; see the warning above). - 2026-10-02, ADR-1466. The SYCL twin adds every per-scale sum in the CPU's raster order; it is exact by construction.
- ADR-1414. The SYCL twin computes the CPU's arithmetic in exact float pairs. The HIP twin computes the CPU's arithmetic as well (ADR-1403).
- ADR-1221. GPU
clip_dbbecame a ceiling on the dB output (see the warning under Options). - ADR-1299. The SYCL twin gained
enable_chromaand the CPU extractor checks every scored plane against the minimum size.
See also¶
- SSIM - single-scale structural similarity
- SSIMULACRA2 - perceptually tuned alternative