Research 0058 — canonical-6 permutation importance for vmaf_tiny_v2¶
Research-0058¶
- Status: Complete
- Date: 2026-05-03
- Author: Lusoris (Claude Opus 4.7 agent)
- Companion ADRs: none — empirical research only, no decision change
- Companion code:
scripts/dev/permutation_importance.py
Question¶
Of the canonical-6 features fed to the shipped vmaf_tiny_v2.onnx regressor (adm2, vif_scale0..3, motion2), which carry the strongest signal toward the teacher VMAF prediction, and are any near-zero (candidates for removal in a future v3+)?
Method¶
Permutation importance, model-agnostic and training-free:
- Load
model/tiny/vmaf_tiny_v2.onnxvia onnxruntime CPU EP. The model has the StandardScaler baked in as Constant nodes (pervmaf_tiny_v2.json), so raw feature values are fed directly. - Sample 5000 rows uniformly at random (seed 20260503) from
runs/full_features_4corpus.parquet(330 499 frames over Netflix Public, KoNViD, BVI-A, BVI-B, BVI-C, BVI-D). - Compute baseline PLCC of the ONNX prediction vs the teacher
vmafcolumn. - For each feature column, permute its values (
np.random.permutation) and re-score. Repeat 5 times with seeds1000..1004. Report mean ± std. - The drop in PLCC is the feature's importance.
Wall time: ~3 seconds on CPU (no GPU needed). Reproducible via the script above.
Results¶
Baseline PLCC: 0.999897 (5000-row sample).
| Rank | Feature | Baseline PLCC | After permutation | Drop ± std |
|---|---|---|---|---|
| 1 | adm2 | 0.9999 | 0.5470 | +0.4529 ± 0.0040 |
| 2 | motion2 | 0.9999 | 0.6545 | +0.3454 ± 0.0037 |
| 3 | vif_scale3 | 0.9999 | 0.8950 | +0.1049 ± 0.0007 |
| 4 | vif_scale2 | 0.9999 | 0.9439 | +0.0560 ± 0.0004 |
| 5 | vif_scale1 | 0.9999 | 0.9968 | +0.0031 ± 0.0001 |
| 6 | vif_scale0 | 0.9999 | 0.9981 | +0.0018 ± 0.0001 |
Findings¶
adm2is the dominant signal (drop −0.453). Permuting it collapses PLCC from 0.9999 to 0.55 — the model effectively cannot reconstruct VMAF without ADM. This matches Netflix's own ranking in the original VMAF paper (ADM2 + VIF + motion).motion2is a strong second (drop −0.345). Temporal energy contributes meaningfully despite the scalar-per-frame nature. Worth keeping; cannot be substituted by spatial features.vif_scale3andvif_scale2carry the bulk of the VIF signal (drops −0.105 and −0.056). The coarse VIF scales encode global luminance fidelity that ADM (edge-focused) misses.vif_scale0andvif_scale1are near-zero (drops +0.0018 and +0.0031, both <1 % PLCC each). The fine VIF scales are essentially redundant withvif_scale2/3in the small-MLP regime — the model has learned to weight them ~0. Consistent with classic VIF's known high-scale dominance for natural-content distortions.
Decision implication¶
A future vmaf_tiny_v3 (or a tinier vmaf_nano) experiment dropping vif_scale0 and vif_scale1 from the input vector should be tractable — expected PLCC delta on the held-out set is on the order of 0.005, which may be acceptable for a 4-feature canonical input that halves the VIF extractor work. Not making that decision here; flagging as a follow-up hypothesis for the canonical-N exploration backlog.
This research digest does not mandate any code or model change. The shipped vmaf_tiny_v2 keeps its 6-feature input.
Reproducer¶
Produces the table above in ~3 seconds against the shipped ONNX and the 4-corpus parquet in runs/.
ADR-0108 deliverables (research-only PR)¶
- Research digest — this file.
- Decision matrix — no alternatives: only-one-method study (permutation importance is the standard model-agnostic technique; no training-based alternative was within the time-box).
AGENTS.mdinvariant note — no rebase-sensitive invariants. This PR adds one researcher script and one doc; nothing patches upstream surfaces.- Reproducer / smoke-test command — see "Reproducer" above.
CHANGELOG.mdentry — opt-out: research-only, no user-visible change.docs/rebase-notes.mdentry — no rebase impact: research-only doc ev script underscripts/dev/.
References¶
model/tiny/vmaf_tiny_v2.json— sidecar with baked-in scaler stats.- PR #250 —
vmaf_tiny_v2ship commit3999cdab. - docs/research/0046-vmaf-tiny-v3-mlp-medium-evaluation.md — prior canonical-6 evaluation in the v3 (mlp_medium) regime.