MobileSal saliency (legacy placeholder checkpoint)¶
vmaf_tiny_mobilesal_placeholder_v0 — the historical smoke checkpoint for the no-reference saliency feature extractor (mobilesal). The extractor runs a tiny ONNX saliency model over the distorted frame and emits the mean of its per-pixel saliency map as a scalar feature named saliency_mean. It is the scoring-side surface for Wave 1 §2.3 of the tiny-AI roadmap, the first half of backlog item T6-2 (T6-2a). Encoder-side ROI tooling (tools/vmaf-roi, per-CTU QP-offset sidecars) is shipped as T6-2b.
Legacy smoke placeholder: not for production
model/tiny/mobilesal.onnx is a synthetic smoke placeholder (3 to 1 channel Conv + Sigmoid). It matches the MobileSal I/O contract to validate pipeline wiring, but emits an almost constant saliency (about 0.5).
For production saliency use saliency_student_v2, the fork-trained default since 2026-05-15 (ADR-0444), or saliency_student_v1 (ADR-0286). Both keep the same input / saliency_map tensor names and run through the same feature_mobilesal.c extractor.
Upstream MobileSal weights remain deferred by ADR-0257 (CC BY-NC-SA 4.0, Google-Drive-walled, RGB-D); the fork-trained students provide the license-clean production path.
Upstream paper: Wu, Liu, Cheng, Lu, Cheng, "MobileSal: Extremely Efficient RGB-D Salient Object Detection", IEEE TPAMI 2021.
What the output means¶
The extractor emits a single feature named saliency_mean, one value per frame (the distorted frame is the input — saliency is no-reference).
| Value | Interpretation |
|---|---|
| ~0.0 | Flat / featureless content; no salient subject |
| ~0.2 – 0.4 | Typical natural-content frame |
| ~0.5 | Foreground subject occupies a sizeable fraction of the frame |
| ~0.8+ | Subject dominates; mostly-salient content |
| 1.0 | Saturated (every pixel maxed) — usually a sign of model misuse |
Saliency-mean is not a quality score on its own — it is a content descriptor. Downstream consumers correlate saliency_mean against existing metric features (e.g. vmaf, lpips, psnr) to study how foreground-vs-background distortion affects subjective quality.
The full saliency map is computed internally to derive the mean; it is intentionally not exposed as a per-pixel feature in T6-2a. The encoder side (tools/vmaf-roi per-CTU QP-offset sidecar) consumes the same model in T6-2b and exports the map in encoder-native format.
Shipped checkpoint¶
| Field | Value |
|---|---|
| Model id | mobilesal_placeholder_v0 |
| Display name | vmaf_tiny_mobilesal_placeholder_v0 |
| Location | model/tiny/mobilesal.onnx |
| Size | 330 bytes (synthetic placeholder) |
| SHA-256 | f122631089977c4be7d60b9bf3d4daf186d275bd0587db2c9878578e006b91d4 |
| ONNX opset | 17 |
| Upstream source (paper) | yuhuan-wu/MobileSal (HEAD 8f42ded5; not currently shippable — see ADR-0257) |
| License (placeholder) | BSD-2-Clause-Patent (this fork) |
| License (upstream MobileSal weights) | CC BY-NC-SA 4.0 — incompatible with the fork; per yuhuan-wu/MobileSal/README.md §License. ADR-0218's MIT claim was inaccurate; corrected here and in ADR-0257. |
| Exporter (placeholder) | ai/scripts/gen_mobilesal_placeholder_onnx.py |
| Registry entry | mobilesal_placeholder_v0 in model/tiny/registry.json (smoke=true) |
| Status | Legacy smoke placeholder — superseded for production by saliency_student_v2 (ADR-0444) / saliency_student_v1 (ADR-0286) |
The placeholder ONNX is deterministic (no doc_string, fixed producer_version, deterministic protobuf serialisation) so the sha256 stays stable across re-runs of the export script.
For content-dependent saliency, point the extractor at the production default model/tiny/saliency_student_v2.onnx (or model/tiny/saliency_student_v1.onnx). The placeholder is retained to keep the historical ABI / I/O-contract smoke path available.
Input / output contract¶
The C extractor binds tensors by name, so any future drop-in (real upstream MobileSal export, distilled student, etc.) must declare the exact same names:
inputs:
input float32[1, 3, H, W] ImageNet-normalised RGB, NCHW
outputs:
saliency_map float32[1, 1, H, W] per-pixel saliency in [0, 1]
H and W are dynamic — both the placeholder and the upstream graph match whatever resolution the C side feeds. ImageNet normalisation (mean [0.485, 0.456, 0.406], std [0.229, 0.224, 0.225]) is applied in the C side via the shared vmaf_tensor_from_rgb_imagenet() helper, identical to LPIPS's wiring (see lpips_sq_v1.md).
Usage — CLI¶
vmaf \
--reference ref.yuv \
--distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--feature mobilesal=model_path=model/tiny/saliency_student_v2.onnx \
--output score.json
The --feature argument takes name=option=value (colon-separated for more options); there is no --feature_params option. The output JSON gains a per-frame saliency_mean column alongside any other features requested in the same run. Combine with lpips and vmaf for the full saliency-quality picture:
vmaf --reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--feature vmaf \
--feature lpips=model_path=model/tiny/lpips_sq.onnx \
--feature mobilesal=model_path=model/tiny/saliency_student_v2.onnx \
--output combined.json
Equivalently, set the model path via env var:
VMAF_MOBILESAL_MODEL_PATH=model/tiny/saliency_student_v2.onnx \
vmaf --reference ref.yuv --distorted dist.yuv \
--width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
--feature mobilesal --output score.json
Usage — C API¶
#include <libvmaf/libvmaf.h>
VmafFeatureDictionary *opts = NULL;
vmaf_feature_dictionary_set(&opts, "model_path", "model/tiny/saliency_student_v2.onnx");
int err = vmaf_use_feature(ctx, "mobilesal", opts);
/* ... vmaf_score_pooled(ctx, ..., "saliency_mean", ...) for the per-frame mean */
Equivalent to setting VMAF_MOBILESAL_MODEL_PATH before vmaf_use_feature(ctx, "mobilesal", NULL).
Known limitations¶
| Limit | Behaviour | Workaround |
|---|---|---|
| Bit depth | 8-bit YUV only; other depths are rejected at init() with -ENOTSUP (see the message below) | Drop --feature mobilesal from HDR / 10-bit / 12-bit runs, or use --bitdepth 8 when the source is 8-bit content in a 10-bit container |
| Pixel format | YUV420P, YUV422P, YUV444P accepted; YUV400P (luma-only) rejected at init() because the model needs three RGB channels | Convert to a chroma-carrying format |
| Colour space | BT.709 limited-range Y'CbCr to RGB on the C side, matching feature_lpips.c; BT.2020 / full range is approximate (deliberate trade-off, see the feature_mobilesal.c comment) | None |
| Resolution | Bounded by the selected ONNX graph's dynamic shape; the placeholder has no useful quality floor, the fork-trained student was trained on 256x256 crops. The student checkpoints need sides that are multiples of 8; the extractor pads other sizes (576x324 runs as 576x328, the last column and row repeated) and averages the map over the frame's own area (ADR-1540) | None for the extractor; pad yourself when running the ONNX graph directly |
| Execution provider | The extractor opens its session with the default auto device: CUDA, OpenVINO GPU, ROCm, CoreML, then CPU, whichever the linked ONNX Runtime provides; independent of -Denable_cuda and of --tiny-device | None needed |
| Score interpretation | With the placeholder, saliency_mean is about 0.5 whatever the input; the placeholder only locks down the pipeline. With a student checkpoint the score is content-dependent | Use a student checkpoint |
With --bitdepth 10 or 12 together with --feature mobilesal the run aborts before scoring. The saliency model requires 8-bit ImageNet-normalised RGB; wider depths would require retraining and are not planned (ADR-0613 §P1-3). The message reads:
mobilesal: bpc=10 is not supported (8-bit only). The mobilesal extractor
requires 8-bit YUV input because the saliency model was trained on 8-bit
ImageNet-RGB. Use --bitdepth 8 or omit --feature mobilesal for HDR /
10-bit / 12-bit content.
Training and evaluation¶
The placeholder is not trained: it is a generated Conv + Sigmoid with fixed weights, so no training data, quality evaluation or correlation figure exists for it. Trained, evaluated saliency weights are on the saliency_student_v1 and saliency_student_v2 cards.
How the placeholder is regenerated¶
python ai/scripts/gen_mobilesal_placeholder_onnx.py # rewrite model/tiny/mobilesal.onnx
python ai/scripts/gen_mobilesal_placeholder_onnx.py --check # exit 1 if the file differs
# wrote model/tiny/mobilesal.onnx (330 bytes, sha256 f12263...)
The generator writes the ONNX file only; the sidecar JSON and the registry entry are committed files. The output is byte-identical to the shipped file, and --check fails when it is not. The sha256 in registry.json is verified before CreateSession.
Related¶
lpips_sq_v1.md— sister full-reference DNN extractor; shares the YUV → ImageNet-RGB plumbing.../roadmap.md§2.3 — Wave 1 MobileSal scope.saliency_student_v2.md— production default saliency weights for this extractor (ADR-0444).saliency_student_v1.md— initial fork-trained saliency student baseline (ADR-0286).- ADR-0218 — design notes (smoke-only placeholder, scoring-vs-encoder split, scalar-vs-map output).
- ADR-0257 — first blocker (T6-2a-followup real-weights swap deferred): upstream MobileSal license, distribution and RGB-D mismatch.
- Research-0053 — upstream survey, licence analysis, and alternatives walk.
- ADR-0265 — second blocker: U-2-Net
u2netpdistribution + op-allowlist mismatch. - Research-0055 — companion survey for ADR-0265.
- ADR-0286 — fork-trained production saliency-student path.
- ADR-0042 — tiny-AI doc-substance rule this page satisfies.