SSIMULACRA 2 Extractor¶
SSIMULACRA 2 is a full-reference perceptual similarity metric from the JPEG XL ecosystem. It compares reference and distorted frames in an XYB-inspired colour space, combines multi-scale SSIM-style structural terms, and applies an asymmetric penalty for lost texture energy. The fork ships it as a normal libvmaf feature extractor named ssimulacra2.
Output¶
| Field | Value |
|---|---|
| Feature name | ssimulacra2 |
| Output metric | ssimulacra2 |
| Direction | Higher is better |
| Range | [0, 100]; identical frames return 100 |
| Snapshot gate | python/test/ssimulacra2_test.py |
Practical score bands:
| Score | Meaning |
|---|---|
90-100 | Visually lossless |
70-90 | High quality |
50-70 | Medium quality, clearly lossy |
30-50 | Low quality |
0-30 | Very low quality |
Usage¶
vmaf \
--reference ref.yuv \
--distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--feature ssimulacra2 \
--output score.json
The per-frame JSON metric key is ssimulacra2; pooled values appear under the same key in pooled_metrics.
Options¶
| Option | Type | Default | Values | Effect |
|---|---|---|---|---|
yuv_matrix | int | 0 | 0..3 | 0: BT.709 limited, 1: BT.601 limited, 2: BT.709 full, 3: BT.601 full |
Example:
Inputs¶
- Pixel formats: YUV 4:2:0, 4:2:2, and 4:4:4. The colour conversion needs both chroma planes, so 4:0:0 (luma-only) input is refused at init with
ssimulacra2: needs a YUV 4:2:0, 4:2:2 or 4:4:4 input, not 4:0:0and-EINVALfromvmaf_read_pictures(). The CUDA, SYCL, HIP and Metal twins refuse it the same way, with their own name in the message (ssimulacra2_metal: needs ...). When a model or--backendpicked the CUDA, SYCL or HIP twin, 4:0:0 input goes to the CPU extractor first, which then refuses it. ThevmafCLI never produces 4:0:0 input: it rejects-p 400and converts Y4Mmonoinput to 4:2:0. - Bit depths: 8, 10, and 12 bpc.
- CPU SIMD: AVX2, AVX-512, NEON, and SVE2 when the host advertises it. The CPU scalar/SIMD path is bit-exact across the fork's host matrix.
GPU twins¶
The CUDA, SYCL and HIP twins run the whole frame on the device and return the CPU extractor's score bit for bit. The Metal twin offloads the pyramid blur and per-pixel multiply stages while keeping the colour conversion, XYB, downsample and final combine on the host. (The Vulkan backend was removed in ADR-0726.)
| Twin | Invocation | Exact vs CPU | ADR | Evidence (fragment) | Measured on |
|---|---|---|---|---|---|
ssimulacra2_cuda | --backend cuda | Yes | ADR-1391, ADR-1433 | scripts/ci/exact_twins.d/ssimulacra2.cuda | RTX 4090: 113 of 113 frames |
ssimulacra2_sycl | --backend sycl | Yes | ADR-1363, ADR-1446 | scripts/ci/exact_twins.d/ssimulacra2.sycl | Arc A380 (xe): 266 of 266 frames |
ssimulacra2_hip | --backend hip | Yes | ADR-1390, ADR-1445 | scripts/ci/exact_twins.d/ssimulacra2.hip | gfx1036 (ROCm 7.2.4): 178 of 178 frames |
ssimulacra2_metal | --backend metal | No (FEATURE_TOLERANCE, 5e-3) | none | none | not recorded |
The measurements are at --precision max against --backend cpu, and with every yuv_matrix:
- CUDA: Netflix 576x324 at 8, 10, 12 and 16 bits, both 1080p checkerboard pairs, BBB 3840x2160.
- SYCL: Netflix 576x324 at 8, 10, 12 and 16 bits and as 10-bit 4:2:2, both 1080p checkerboard pairs, 200 frames of BBB 3840x2160.
- HIP: Netflix 576x324 at 8, 10, 12 and 16 bits and as 10-bit 4:2:2, both 1080p checkerboard pairs, Sparks 480x270, full-range noise at four bit depths, a bright 16-bit 1080p pair, 48 frames of BBB 3840x2160.
Speed and memory¶
| Twin | Frame | Time now | Time before exact sums | CPU extractor | Memory |
|---|---|---|---|---|---|
| CUDA, RTX 4090 | 3840x2160 | 15.6 ms | 7.8 ms (ADR-1433) | 126 ms on sixteen threads | About 1.4 GB device (14.5 full-size three-plane float buffers: the linear-RGB pyramid, XYB, and the two passes of the five blurs); the sums add 3.8 MB |
| SYCL, Arc A380 | 3840x2160 | 195 ms | 84 ms (ADR-1446) | About 125 ms on sixteen threads | not recorded |
| SYCL, Arc A380 | 576x324 | 13.5 ms | 5.3 ms | 1.4 ms on sixteen threads | not recorded |
| HIP, gfx1036 | 3840x2160 | 662 ms | 234 ms (ADR-1445) | 124 ms on sixteen threads | not recorded |
| HIP, gfx1036 | 1920x1080 | 167 ms | 58 ms | not recorded | not recorded |
The CPU is faster on integrated GPUs and on the Arc A380
On the A380 the CPU is the faster path for this metric at both sizes (before ADR-1446 the twin was faster at 3840x2160). On the integrated gfx1036 the CPU is the faster path too. Most of the SYCL increase is the integer arithmetic of the terms; the pass over the chunks runs on one lane per sum and is what small frames pay (T-SYCL-SSIMULACRA2-EXACT-THROUGHPUT-2026-10-02). On HIP most of the increase is the double-precision terms, evaluated twice per scale, and the integer steps (T-HIP-SSIMULACRA2-EXACT-THROUGHPUT-2026-10-02).
How the exact sums work¶
The CPU evaluates six terms per pixel and channel in double precision and adds each into one double, pixel after pixel, and every such add rounds. The twins return the result of that loop without running it on one thread.
- While the running sum stays between two powers of two, adding a term moves it by a whole number of steps. The device adds those whole numbers per chunk of pixels in parallel (1024 pixels on CUDA and HIP, 512 on SYCL), and one pass over the chunks puts them together.
- The few chunks in which the sum passes a power of two are added term by term.
- CUDA and HIP devices have double precision, so their per-pixel terms are the CPU's own double expressions. A SYCL device has no double precision, so the SYCL twin computes each term's double in 64-bit integers (a sign, a 53-bit significand and an exponent), the CPU's operations one for one.
CUDA: device-resident, one readback per frame¶
ssimulacra2_cuda (ADR-1391) reads the Y/U/V planes that the CUDA picture pipeline uploads once for every CUDA extractor, so the twin makes no copy of its own. Its only transfer is one 864-byte block of per-scale sums per frame. Name it, or pass --backend cuda --feature ssimulacra2, which maps to the twin for inputs it can run and to the CPU extractor otherwise (4:0:0 input and frames below 8x8):
vmaf -r ref.yuv -d dis.yuv -w 3840 -h 2160 -p 420 -b 8 \
--backend cuda --no_prediction --feature ssimulacra2_cuda -o out.json --json
Check and time it with (--vmaf takes an absolute path; --bbb-dir is explained below):
python3 scripts/dev/speed_gpu_parity.py --backend cuda --feature ssimulacra2 \
--vmaf "$PWD/build/tools/vmaf" \
--netflix-dir python/test/resource/yuv --bbb-dir testdata/bbb
SYCL: device-resident, one readback per frame¶
ssimulacra2_sycl (ADR-1363) uploads the raw Y/U/V planes once, converts to linear RGB and XYB, blurs, forms the SSIM and edge-difference sums and downsamples on the device, and reads back one 864-byte block of per-scale sums. Name it, or pass --backend sycl --feature ssimulacra2, which maps to the twin for inputs it can run and to the CPU extractor otherwise (ADR-1359; the twin needs chroma planes and at least 8x8, and rejects 4:0:0 input at init):
vmaf -r ref.yuv -d dis.yuv -w 3840 -h 2160 -p 420 -b 8 \
--backend sycl --no_prediction --feature ssimulacra2_sycl -o out.json --json
Check and time it with:
ONEAPI_DEVICE_SELECTOR=level_zero:0 python3 scripts/dev/speed_gpu_parity.py \
--backend sycl --feature ssimulacra2 --vmaf "$PWD/build/tools/vmaf" \
--netflix-dir python/test/resource/yuv --bbb-dir testdata/bbb
HIP: device-resident, tiled row pass¶
ssimulacra2_hip (ADR-1390) uploads the raw Y/U/V planes once into pinned staging in submit(), converts to linear RGB and XYB, executes IIR blurs with a tiled shared-memory row pass (SS2H_ROW_TILE rows per wavefront, single-wave blocks, two-slot ring, register prefetch), evaluates the per-pixel SSIM and edge terms in double precision, adds them with the result of the CPU's loops, and downsamples on the device. A single 864-byte readback of per-scale sums occurs in collect(). Name it, or pass --backend hip --feature ssimulacra2 (ADR-1359). The twin rejects 4:0:0 input at init:
vmaf -r ref.yuv -d dis.yuv -w 3840 -h 2160 -p 420 -b 8 \
--backend hip --no_prediction --feature ssimulacra2_hip -o out.json --json
Check and time it with:
python3 scripts/dev/speed_gpu_parity.py --backend hip --feature ssimulacra2 \
--vmaf "$PWD/build-hip/tools/vmaf" \
--netflix-dir python/test/resource/yuv --bbb-dir testdata/bbb
Checking a twin¶
scripts/dev/speed_gpu_parity.py compares the twin with the CPU extractor frame by frame. Its default bound is 0: every frame must be bit-identical.
The script always runs two fixtures: the Netflix 576x324 pair from --netflix-dir, and a Big Buck Bunny 3840x2160 8-bit 4:2:0 pair that you supply as ref_3840x2160_200f.yuv and dis_3840x2160_200f.yuv in --bbb-dir. The default testdata/bbb is not tracked in the repository, so pass --bbb-dir explicitly.
Implementation notes: cross-compiler bit-exactness (FMA unification)¶
picture_to_linear_rgb performs the YCbCr to linear-RGB conversion using a fused-multiply-add (FMA) chain on every code path: scalar uses fmaf() and the AVX2 / AVX-512 / NEON main loops use _mm256_fmadd_ps / _mm512_fmadd_ps / vfmaq_f32. This preserves the left-to-right associativity of G = Yn + cb_g*Un + cr_g*Vn (two chained FMAs) while delivering a single, identically-rounded result on every supported compiler. See ADR-0891 for the analysis and alternatives.
The pipeline is ill-conditioned downstream, which matters when reading any ssimulacra2 number:
- The edge-diff term computes
|img - blur(img)|, a catastrophic cancellation. - Pooling takes a 4-norm, which is dominated by the few largest survivors.
- A difference of about 1 ULP in linear RGB therefore grew to a 2.62e-03 difference in the final score, not the about 1e-7 a single rounding step would suggest.
After ADR-1205 every shipped path uses the same FMA chain and CPU-vs-CUDA agreement is about 2.8e-09.
If you add another ssimulacra2 code path, the conversion must be fmaf()-based and in the ADR-0891 order. core/test/test_ssimulacra2_simd.c compares the SIMD kernels against a private scalar reference rather than the shipped function, so it will not catch a shipped copy that drifts.
Limitations¶
- Chroma is nearest-neighbour upsampled to luma resolution before colour conversion.
- The recursive Gaussian coefficient derivation is pinned to the shipped sigma used by the libjxl reference path; arbitrary sigma values are not a user option.
- Scores are useful as a perceptual ranking signal, not as an MOS-calibrated VMAF replacement.
History¶
- ADR-1445 (HIP), ADR-1446 (SYCL), ADR-1433 (CUDA). Before them the terms were added in a fixed tree (pairs of floats on SYCL and HIP) and no SYCL or HIP frame was identical. The score was up to 7.6e-11 from the CPU's on SYCL and HIP and up to 7.3e-11 on CUDA;
ssimulacra2_sycl,ssimulacra2_hipandssimulacra2_cudascores stored before differ from new ones by that much. - ADR-1391, ADR-1363, ADR-1390. The CUDA, SYCL and HIP twins became device-resident. Before ADR-1391 the CUDA twin copied every scale between device and host; before ADR-1363 the SYCL twin copied about 4 GB per 4K frame between device and host.
- ADR-1205.
picture_to_linear_rgb's FMA chain reached the five shipped copies that ADR-0891 had not: the scalar fallback incore/src/feature/ssimulacra2.cand the host-side conversions in the CUDA, HIP, Metal and SYCL twins. Those kept a plainmul + add, so until ADR-1205 a host without AVX2 scoredssimulacra2differently from one with it, and every GPU backend disagreed with the CPU. - ADR-0891. Earlier revisions used explicit
mul + addpairs, but icx with-mfmaauto-fused them while gcc kept them separately-rounded, producing sub-ULP divergence between compilers. It reached the four SIMD kernels and the SIMD test's own scalar reference.