SSIM¶
SSIM (Structural Similarity Index Measure) quantifies perceptual image quality by comparing luminance, contrast, and structure between a reference and a distorted frame. The fork's integer SSIM extractor scores the luma plane, uses fixed-point arithmetic compatible with the upstream Netflix reference, and feeds the ssim feature that VMAF model JSON files list.
How to run¶
The extractor is registered as ssim. The CLI also accepts integer_ssim as an alias for it; the C API (vmaf_use_feature(ctx, "integer_ssim")) does not.
# SSIM (default, linear)
build/tools/vmaf \
--reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--no_prediction --feature ssim --output /dev/stdout
# SSIM in dB, capped for identical frames
build/tools/vmaf \
--reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--no_prediction --feature ssim=enable_db=true:clip_db=true \
--output /dev/stdout
# The same on the SYCL twin (integer_ssim_sycl)
build/tools/vmaf \
--reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--backend sycl --no_prediction \
--feature ssim=enable_db=true:clip_db=true --output /dev/stdout
Options follow the feature name after the first =, separated by :.
Output features¶
| Feature name | Description | Condition |
|---|---|---|
ssim | SSIM of the luma (Y) plane, or its dB form with enable_db | Always |
The extractor is luma-only; it has no chroma option.
Options¶
| Option | Type | Default | Effect |
|---|---|---|---|
enable_db | bool | false | Report \(-10 \log_{10}(1 - \mathrm{ssim})\) instead of the linear score. A perfect score (identical frames) is +inf. |
clip_db | bool | false | Cap the dB value at \(\left\lceil 10 \log_{10} \dfrac{\mathrm{peak}^2}{0.5 / (w\,h)} \right\rceil\), the dB of half a sample of error over the frame. Needs enable_db to have an effect. |
The CPU extractor and the CUDA, HIP, SYCL and Metal twins accept both options. A model that sets an option the active backend's twin lacks computes ssim on the CPU (ADR-1183).
Variants¶
| Registered name | Backend | Algorithm | Precision vs CPU |
|---|---|---|---|
ssim | CPU | Integer fixed-point | Reference |
integer_ssim_cuda | CUDA | Real int64 moments + double SSIM, summed in the CPU's order | Bit-identical at every frame size |
integer_ssim_hip | HIP | Real int64 moments + double SSIM, summed in the CPU's order | Bit-identical at every frame size |
integer_ssim_sycl | SYCL | Real int64 moments + the CPU's double SSIM term computed in 64-bit integers, summed in the CPU's order | Bit-identical at every frame size and bit depth |
integer_ssim_metal | Metal | Fixed-point, two-pass separable Gaussian | places=4 (target, ADR-0214) |
All twins publish the feature name ssim, so existing VMAF model JSON files work unchanged. Their sources are ssim_cuda.c (CUDA), integer_ssim_hip.c (HIP), integer_ssim_sycl.cpp (SYCL) and integer_ssim_metal.mm + integer_ssim.metal (Metal, ADR-0564 cross-backend completion).
Selecting a twin¶
- By model. A model that lists the
ssimfeature gets the twin of the active backend (--backend cuda|hip|sycl|metal) automatically. - By feature name.
--backend cuda|hip|sycl|metal --feature ssimruns the same twin when it can honour the options you set, and the CPU extractor otherwise (ADR-1359). - By twin name. Naming a twin, for example
--backend hip --feature integer_ssim_hip, always registers that twin.
Note
There is also a float_ssim variant on CUDA (11-tap floating-point Gaussian, a distinct algorithm). Before ADR-0564 the CUDA and HIP backends silently returned float_ssim scores in the "ssim" field. This bug is fixed.
GPU twins¶
The CUDA, HIP and SYCL twins give the CPU's value for every frame, down to the last bit of the --precision max output, with and without enable_db and clip_db. Each computes every per-pixel term as the CPU does, reads the terms back and adds them on the host in the CPU's order. They report the CPU's value exactly, finite or +inf, and apply the options on the host to the device-reduced score (ADR-1365, ADR-1373, ADR-1382; the HIP twin's gfx1036 run is in T-BUG048-GPU-OPTION-PARITY-REMAINDER-2026-09-26 in state.md).
| Twin | Exact vs CPU | ADR | Evidence (fragment) | Cost at 3840x2160 | Extra memory at 3840x2160 |
|---|---|---|---|---|---|
integer_ssim_cuda | Yes | ADR-1424 | scripts/ci/exact_twins.d/ssim.cuda | 9.7 ms instead of 2.2 ms (RTX 4090); not measurable at 576x324 | 66 MB device, as much pinned host |
integer_ssim_hip | Yes | ADR-1438 | scripts/ci/exact_twins.d/ssim.hip | 98.1 ms instead of 94.3 ms (gfx1036); 30.0 ms instead of 28.2 ms at 1920x1080 | 66 MB device, as much pinned host |
integer_ssim_sycl | Yes | ADR-1443 | scripts/ci/exact_twins.d/ssim.sycl | 31.9 ms instead of 17.8 ms (Arc A380; 0.78 ms instead of 0.45 ms at 576x324); the CPU extractor takes 111 ms | 66 MB pinned host; device memory unchanged |
integer_ssim_metal | Tolerance (FEATURE_TOLERANCE, ADR-0214) | ADR-0564 | none (not declared exact) | no figure recorded | no figure recorded |
The CPU's sum is sequential, so an exact twin pays for reading the terms back: that is the cost column. At large frames the twin is slower than a device reduction would be.
What exactness means when you use it¶
--backend cuda --feature ssim gives the same number as --backend cpu --feature ssim for every frame, down to the last bit of the --precision max output, with and without enable_db / clip_db (the same holds for --backend hip and --backend sycl).
- You can mix CPU and twin
ssimresults in one data set. - Stored outputs from before the exact-sum ADRs differ from new ones: up to 1.1e-11 (3.6e-10 in dB) for CUDA, up to 1.1e-11 on frames above 4096 pixels for HIP, and 7e-9 to 3e-7 on video for SYCL.
Measured coverage¶
| Twin | Device | Measured on |
|---|---|---|
| CUDA | RTX 4090 | Netflix 576x324 pair at 8, 10, 12 and 16 bits, both 1920x1080 checkerboard pairs, 3840x2160, frames down to 1x1 |
| HIP | gfx1036 | Netflix 576x324 pair at 8, 10, 12 and 16 bits and as 10-bit 4:2:2, both 1920x1080 checkerboard pairs, 48 frames of 3840x2160, full-range noise at four depths, frames down to 1x1 |
| SYCL | Arc A380 | Netflix 576x324 pair at 8, 10, 12 and 16 bits and 4:2:2, both 1920x1080 checkerboard pairs, 200 frames of 3840x2160, frames down to 1x1 |
SYCL specifics¶
A SYCL kernel has no double on the GPUs the backend targets. The twin therefore runs the CPU's double-precision operations for each pixel in 64-bit integers, which gives the CPU's double exactly, and adds the per-pixel values on the host in the CPU's order.
- 16-bit input works. Before ADR-1443 the twin's single-precision term overflowed on 16-bit frames and the run stopped with
invalid ratio. - dB values. With
enable_db, the linear score is the CPU's and the dB value is computed from it on the host. It equals the CPU's when both come from the same build. Between a build made with the Intel compiler and one made with GCC the last digit of the dB value can differ (3.6e-15 measured), because the two link differentlog10implementations; that affects the CPU extractor in the same way.
History¶
- 2026-09-18, HIP twin enters dispatch. It was kept out of dispatch because its kernel was an 11-tap float Gaussian 4.5e-3 away from the CPU. It now runs the CPU's 9-tap int64 kernel and adds the per-pixel terms in the CPU's order, so its score is the CPU's bit for bit; see the HIP backend page.
- ADR-1443, SYCL twin exact. Before it the twin computed in single precision and was 7e-9 to 3e-7 from the CPU on video.
- ADR-1438, HIP twin exact. Before it, frames above 4096 pixels differed from the CPU in the last digits (up to 1.1e-11).
- ADR-1424, CUDA twin exact. Before it the twin differed from the CPU in the last digits (up to 1.1e-11, and 3.6e-10 in dB).
- ADR-0564. Before it, the CUDA and HIP backends returned
float_ssimscores in the"ssim"field.
See also¶
- MS-SSIM - multi-scale structural similarity
- SSIMULACRA2 - perceptually tuned alternative