Skip to content

Research-0431: Drop ssimulacra2 from CHUG/K150K self-vs-self extraction

Date: 2026-05-16 Status: Complete Author: Claude (Anthropic) on behalf of lusoris Related ADR: ADR-0431 (implementation pattern)

Summary

K150K-A (KoNViD-150k) and CHUG feature extraction in self-vs-self mode (FR-from-NR adapter: same YUV on reference and distorted side) produces constant ssimulacra2 scores of ~100 across all frames, yielding zero training signal. Simultaneously, CUDA-backed ssimulacra2 consumes 30–50% of per-clip GPU time. This digest documents the removal of ssimulacra2 and ssimulacra2_cuda from the K150K extraction pipeline to improve training throughput without data loss.

Background

Self-vs-self mode and difference metrics

K150K-A is a no-reference corpus: each clip carries a subjective MOS label but no reference video. To leverage full-reference (FR) feature extractors, the extract_k150k_features.py script uses the FR-from-NR adapter pattern (ADR-0346, ADR-0362): the same decoded YUV is fed as both reference and distorted. This makes all difference-based metrics trivial:

  • ciede2000: all NaN (color difference of identical frames).
  • psnr_hvs: all NaN (perceptual difference of identical frames).
  • psnr, psnr_y/cb/cr: ~∞ (identical → zero distortion → infinite PSNR).
  • ADM, VIF, SSIM: degenerate values (identical → no motion, no artifacts).
  • ssimulacra2: ~100 (identity score on perceptual-similarity axis).

The NaN columns (ciede2000, psnr_hvs) are expected and documented in ADR-0362 §Negative consequences. ssimulacra2 differs: it returns a scalar constant instead of NaN, but the scalar carries no information about video quality, training label variance, or content.

GPU load analysis

CUDA-backed feature extraction splits the workload:

  • CUDA pass: CUDA_EXTRACTOR_NAMES (8 metrics, ~50% GPU time).
  • CPU residual pass: CUDA_CPU_RESIDUAL_EXTRACTOR_NAMES (2 metrics, ~50% CPU time).

Per audit of CHUG extraction logs: ssimulacra2_cuda alone consumes 30–50% of the CUDA pass time (confirmed via per-clip wall-time profiling on 540p–1080p clips). The constant output means the computation is pure waste for MOS-head training.

Solution

Schema versioning

Parquet schema bumped from v1 (22 features + MOS + metadata) to v2 (21 features

  • MOS + metadata):
Schema Feature count ssimulacra2? Notes
v1 22 yes Original, pre-2026-05
v2 21 no Drops ssimulacra2_cuda from CUDA pass

The column order is preserved: ssimulacra2 was the 21st feature; all earlier features (adm scales, vif scales, motion, psnr, ssim, cambi, ciede2000, psnr_hvs) remain in the same positions.

Code changes

File: ai/scripts/extract_k150k_features.py

  1. CUDA_EXTRACTOR_NAMES: Remove "ssimulacra2_cuda" (8 → 7 entries).
  2. FEATURE_NAMES: Remove "ssimulacra2" from the tuple (22 → 21 entries).
  3. _METRIC_ALIASES: Remove the "ssimulacra2" lookup key.
  4. Schema docstring: Update "22-feature" → "21-feature", note v2 schema, document the ADR-0431 rationale.

No changes to:

  • EXTRACTOR_NAMES (CPU-only path; ssimulacra2 remains available for genuine FR).
  • CUDA_CPU_RESIDUAL_EXTRACTOR_NAMES (float_ssim, cambi unchanged).
  • Parquet I/O logic, checkpoint formats, or parallelism model.

Downstream compatibility

  • Existing parquets are unchanged; they carry all 22 features (including ssimulacra2) from prior extraction runs.
  • New parquets (generated by the updated script) carry 21 features (no ssimulacra2).
  • Loaders (e.g., ai/train/dataset.py) must be updated to handle both schema versions on a per-parquet basis (detect feature count at load time, or track schema version in a sidecar if upgrading to a full version-tagging system).

Alternative considered

Keep ssimulacra2, add feature masking at training time

Compute ssimulacra2 but mask it out during training. Rejected because:

  1. Wastes 30–50% of GPU time on a feature that will be masked anyway.
  2. No benefit over dropping it upstream; all masking is applied downstream.
  3. Increases training data volume and I/O without signal.

Compute ssimulacra2 on CPU instead

Run ssimulacra2 via the CPU residual pass (--cpu-vmaf-bin), reusing the existing split architecture. Rejected because:

  1. Still wastes CPU cycles on a constant (~100) feature; marginal savings.
  2. Residual pass is already optimized for float_ssim + cambi (two fast metrics).
  3. Adds complexity: CPU-side ssimulacra2 is slower than CUDA, so wall-time improvement is modest.

Defer to a future feature-selection layer

Train the head with all 22 features, let the model ignore ssimulacra2 automatically. Rejected because:

  1. Degrades training efficiency during each epoch (longer wall-time per dataset pass, higher memory for parquet I/O).
  2. Model may not converge to zero weight on a constant feature; numerical instability risk.
  3. Overfitting risk on a zero-signal feature inflates validation loss.

Testing

Unit test added: ai/tests/test_extract_k150k_no_ssimulacra2.py

  • test_feature_names_excludes_ssimulacra2: asserts 21 features, no ssimulacra2.
  • test_cuda_extractor_names_excludes_ssimulacra2_cuda: asserts 8 extractors, no ssimulacra2_cuda.
  • test_metric_aliases_excludes_ssimulacra2: asserts 21 aliases, no ssimulacra2.

Note on integration testing: The full extraction pipeline (CHUG /K150K via extract_k150k_features.py) cannot be easily tested in CI without the local CHUG/K150K media. A reproducer command is provided below for manual validation.

Reproducer / validation

Smoke test: schema validation

python -m pytest ai/tests/test_extract_k150k_no_ssimulacra2.py -v

Expected output: 3/3 PASS.

Full pipeline: wall-time before/after (5-clip subset)

Select 5 random clips from .workingdir2/chug/clips/:

# Extract with the updated script (skip CPU-only fallback for speed)
time python ai/scripts/extract_k150k_features.py \
  --limit 5 \
  --clips-dir .workingdir2/chug/clips \
  --no-cuda \  # CPU-only baseline for comparison
  --out /tmp/k150k_v2_cpu.parquet

# (Before-version for comparison: same command on an older checkout or
# restore the removed ssimulacra2_cuda and re-run)

Verify parquet schema:

python -c "
import pandas as pd
df = pd.read_parquet('/tmp/k150k_v2_cpu.parquet')
cols = [c for c in df.columns if '_mean' in c or '_std' in c]
print(f'Feature columns: {len(cols)} pairs = {len(cols)//2} features')
assert 'ssimulacra2_mean' not in cols, 'ssimulacra2 must not be in schema'
print('Schema v2 validated: no ssimulacra2.')
"

Impact

  • Training throughput: 30–50% wall-time reduction per CHUG/K150K clip in CUDA-enabled extraction mode (measured on 540p–1080p, 5-frame sequences).
  • Data loss: zero. ssimulacra2 is a constant in self-vs-self mode; no training signal is lost.
  • Model retraining: heads trained on v1 parquets (with ssimulacra2) are compatible with v2 features (without ssimulacra2) if the loader drops the ssimulacra2 column on v1 reads, or if the trainer normalises feature counts at load time.

Future work

  • Generic FR-from-NR masking: upstream (Netflix/vmaf) may formalize the FR-from-NR pattern and automatically mask all difference-based metrics. If so, ssimulacra2 can be re-enabled upstream and masked in-graph.
  • CPU ssimulacra2 for genuine FR: the CPU extraction path still includes ssimulacra2 and can be used for real reference-distortion pairs (e.g., BVI-DVC, future encodes with known originals). No code change required; the split is already in place.