Research-0993: KoNViD / UGC / BVI-DVC Saliency Batch Launch Investigation¶
Date: 2026-06-03 ADR: ADR-0993
Summary¶
Investigation into why 0 saliency materializer runs executed for KoNViD, UGC, and BVI-DVC despite PR #1496 (lusoris/vmaf, now superseded by VMAFx canonical manifests) having merged.
Corpus Status at Investigation Time¶
| Corpus | Path column | Materialisable | Gap |
|---|---|---|---|
| CHUG | src (relative, clips/) | Yes — DONE (5,136 rows) | None |
| Netflix | dis_basename (YUV, dis/) | Yes — DONE (11,190 rows) | None |
| KoNViD-150K | src + width + height | Yes — ready to run | None |
| YouTube UGC | source = corpus ID string | No — path gap | Need file path |
| BVI-DVC D | key = encode params string | No — path gap | Need file path |
Root Cause¶
The UGC and BVI-DVC full-feature parquets were produced by extract_ugc_features.py and bvi_dvc_to_full_features.py respectively. Both scripts emit corpus identifier columns (source, key) rather than absolute file paths. The saliency materializer requires a column that resolves to an actual decodable file.
This is structurally identical to the Netflix issue fixed in PR #540 (where src had to become dis_basename with root=.corpus/netflix/dis/).
KoNViD-150K Resolution¶
konvid_150k.jsonl was generated by konvid_150k_to_corpus_jsonl.py and has:
src: relative mp4 filename (e.g.orig_10000251326_540_5s.mp4)width: 960,height: 540 (all clips)- Clips present at
.corpus/konvid-150k/k150ka_extracted/(152,265 files) - 148,543 unique rows (one per unique clip)
The batch manifest ai/batch-manifests/saliency/konvid-150k.json targets this JSONL directly. No preprocessing required.
UGC Resolution Options¶
The UGC download directory at .corpus/ugc/download/ contains 30 clips per variant (cbr, vod, vodlb, orig), totalling 120 files. Naming convention: {ContentName}_{Variant}.{webm|mp4}. The source column maps as: "ugc-Gaming_1080P-223e-cbr" → Gaming_1080P-223e_cbr.webm.
Option A (recommended): Run youtube_ugc_to_corpus_jsonl.py to produce a path-enriched JSONL, then populate ugc.json:
youtube_ugc_to_corpus_jsonl.py --ugc-dir .corpus/ugc \
--clips-subdir download --clip-suffix _orig.mp4 \
--output .workingdir2/saliency-runs/ugc/ugc_corpus.jsonl
Option B: Re-run extract_ugc_features.py with an additional --emit-src-path flag to write absolute paths alongside the corpus ID.
BVI-DVC Resolution Options¶
Encode keys follow D{ContentName}_{W}x{H}_{fps}fps_{depth}bit_{pix_fmt}. Raw reference YUVs at .corpus/bvi-dvc-raw/ follow A{ContentName}_{W}x{H}_{fps}fps_{depth}bit_{pix_fmt}.yuv (A prefix for reference, D prefix for distorted). All YUVs are 3840×2176.
Option A (recommended): Generate a corpus JSONL for the raw YUVs by stripping the encode-tier letter prefix and deriving the reference YUV path. The default_width=3840 and default_height=2176 fields in bvi-dvc.json are pre-filled.
Option B: Re-run bvi_dvc_to_full_features.py with a path column in output.
Estimated Run Sizes¶
| Corpus | Rows | Unique clips | Estimated run time |
|---|---|---|---|
| KoNViD-150K | 148,543 | 148,543 | ~10–15 h (CPU), ~2–3 h (GPU) |
| YouTube UGC | ~90–120 | ~30 | ~2 minutes |
| BVI-DVC D refs | ~772 | ~193 | ~5 minutes |