Skip to content

Research-0993: KoNViD / UGC / BVI-DVC Saliency Batch Launch Investigation

Date: 2026-06-03 ADR: ADR-0993

Summary

Investigation into why 0 saliency materializer runs executed for KoNViD, UGC, and BVI-DVC despite PR #1496 (lusoris/vmaf, now superseded by VMAFx canonical manifests) having merged.

Corpus Status at Investigation Time

Corpus Path column Materialisable Gap
CHUG src (relative, clips/) Yes — DONE (5,136 rows) None
Netflix dis_basename (YUV, dis/) Yes — DONE (11,190 rows) None
KoNViD-150K src + width + height Yes — ready to run None
YouTube UGC source = corpus ID string No — path gap Need file path
BVI-DVC D key = encode params string No — path gap Need file path

Root Cause

The UGC and BVI-DVC full-feature parquets were produced by extract_ugc_features.py and bvi_dvc_to_full_features.py respectively. Both scripts emit corpus identifier columns (source, key) rather than absolute file paths. The saliency materializer requires a column that resolves to an actual decodable file.

This is structurally identical to the Netflix issue fixed in PR #540 (where src had to become dis_basename with root=.corpus/netflix/dis/).

KoNViD-150K Resolution

konvid_150k.jsonl was generated by konvid_150k_to_corpus_jsonl.py and has:

  • src: relative mp4 filename (e.g. orig_10000251326_540_5s.mp4)
  • width: 960, height: 540 (all clips)
  • Clips present at .corpus/konvid-150k/k150ka_extracted/ (152,265 files)
  • 148,543 unique rows (one per unique clip)

The batch manifest ai/batch-manifests/saliency/konvid-150k.json targets this JSONL directly. No preprocessing required.

UGC Resolution Options

The UGC download directory at .corpus/ugc/download/ contains 30 clips per variant (cbr, vod, vodlb, orig), totalling 120 files. Naming convention: {ContentName}_{Variant}.{webm|mp4}. The source column maps as: "ugc-Gaming_1080P-223e-cbr" → Gaming_1080P-223e_cbr.webm.

Option A (recommended): Run youtube_ugc_to_corpus_jsonl.py to produce a path-enriched JSONL, then populate ugc.json:

youtube_ugc_to_corpus_jsonl.py --ugc-dir .corpus/ugc \
  --clips-subdir download --clip-suffix _orig.mp4 \
  --output .workingdir2/saliency-runs/ugc/ugc_corpus.jsonl

Option B: Re-run extract_ugc_features.py with an additional --emit-src-path flag to write absolute paths alongside the corpus ID.

BVI-DVC Resolution Options

Encode keys follow D{ContentName}_{W}x{H}_{fps}fps_{depth}bit_{pix_fmt}. Raw reference YUVs at .corpus/bvi-dvc-raw/ follow A{ContentName}_{W}x{H}_{fps}fps_{depth}bit_{pix_fmt}.yuv (A prefix for reference, D prefix for distorted). All YUVs are 3840×2176.

Option A (recommended): Generate a corpus JSONL for the raw YUVs by stripping the encode-tier letter prefix and deriving the reference YUV path. The default_width=3840 and default_height=2176 fields in bvi-dvc.json are pre-filled.

Option B: Re-run bvi_dvc_to_full_features.py with a path column in output.

Estimated Run Sizes

Corpus Rows Unique clips Estimated run time
KoNViD-150K 148,543 148,543 ~10–15 h (CPU), ~2–3 h (GPU)
YouTube UGC ~90–120 ~30 ~2 minutes
BVI-DVC D refs ~772 ~193 ~5 minutes