CHUG HDR Audit + Content Splits¶
Finding¶
The CHUG materialisation path already pairs distorted rows with the matching reference content, but it did not make validation leakage hard to do accidentally. CHUG is a bitrate-ladder corpus, so row-level splits can place the same chug_content_name in both train and validation.
The local pipeline also lacked a compact HDR signalling audit. The ingest JSONL records geometry and pix format, but training runs need a pre-flight summary of transfer characteristics, primaries, and malformed PQ/HLG rows before feature rows are trusted.
Change¶
- Add deterministic 80/10/10 content-level splits keyed by
chug_content_name. - Persist
split,chug_split_key, andchug_split_policyinto each CHUG feature row. - Add
--split,--split-seed, and--split-manifestto the feature materialiser. - Add
--audit-outputto write an ffprobe-backed HDR metadata audit with transfer, primaries, pix-fmt, split, missing-file, probe-failure, and malformed-HDR counters. - Teach
train_konvid_mos_head.pyto consume explicit split labels when present, training ontrainrows and validating onval(ortestif no validation rows exist). - Teach the generic
extract_k150k_features.pyFR-from-NR parquet path to preserve CHUG JSONL side metadata via--metadata-jsonl, so local full-feature HDR sweeps do not drop the content-safe split column. - Teach
train_konvid_mos_head.pyto consume those FULL_FEATURES parquet tables directly via--feature-parquet, using<feat>_meanaggregates as canonical trainer features.
Alternatives¶
See ADR-0433.