Skip to content

MobileSal saliency (legacy placeholder checkpoint)

vmaf_tiny_mobilesal_placeholder_v0 — the historical smoke checkpoint for the no-reference saliency feature extractor (mobilesal). The extractor runs a tiny ONNX saliency model over the distorted frame and emits the mean of its per-pixel saliency map as a scalar feature named saliency_mean. It is the scoring-side surface for Wave 1 §2.3 of the tiny-AI roadmap, the first half of backlog item T6-2 (T6-2a). Encoder-side ROI tooling (tools/vmaf-roi, per-CTU QP-offset sidecars) is shipped as T6-2b.

Legacy smoke placeholder: not for production

model/tiny/mobilesal.onnx is a synthetic smoke placeholder (3 to 1 channel Conv + Sigmoid). It matches the MobileSal I/O contract to validate pipeline wiring, but emits an almost constant saliency (about 0.5).

For production saliency use saliency_student_v2, the fork-trained default since 2026-05-15 (ADR-0444), or saliency_student_v1 (ADR-0286). Both keep the same input / saliency_map tensor names and run through the same feature_mobilesal.c extractor.

Upstream MobileSal weights remain deferred by ADR-0257 (CC BY-NC-SA 4.0, Google-Drive-walled, RGB-D); the fork-trained students provide the license-clean production path.

Upstream paper: Wu, Liu, Cheng, Lu, Cheng, "MobileSal: Extremely Efficient RGB-D Salient Object Detection", IEEE TPAMI 2021.

What the output means

The extractor emits a single feature named saliency_mean, one value per frame (the distorted frame is the input — saliency is no-reference).

Value Interpretation
~0.0 Flat / featureless content; no salient subject
~0.2 – 0.4 Typical natural-content frame
~0.5 Foreground subject occupies a sizeable fraction of the frame
~0.8+ Subject dominates; mostly-salient content
1.0 Saturated (every pixel maxed) — usually a sign of model misuse

Saliency-mean is not a quality score on its own — it is a content descriptor. Downstream consumers correlate saliency_mean against existing metric features (e.g. vmaf, lpips, psnr) to study how foreground-vs-background distortion affects subjective quality.

The full saliency map is computed internally to derive the mean; it is intentionally not exposed as a per-pixel feature in T6-2a. The encoder side (tools/vmaf-roi per-CTU QP-offset sidecar) consumes the same model in T6-2b and exports the map in encoder-native format.

Shipped checkpoint

Field Value
Model id mobilesal_placeholder_v0
Display name vmaf_tiny_mobilesal_placeholder_v0
Location model/tiny/mobilesal.onnx
Size 330 bytes (synthetic placeholder)
SHA-256 f122631089977c4be7d60b9bf3d4daf186d275bd0587db2c9878578e006b91d4
ONNX opset 17
Upstream source (paper) yuhuan-wu/MobileSal (HEAD 8f42ded5; not currently shippable — see ADR-0257)
License (placeholder) BSD-2-Clause-Patent (this fork)
License (upstream MobileSal weights) CC BY-NC-SA 4.0 — incompatible with the fork; per yuhuan-wu/MobileSal/README.md §License. ADR-0218's MIT claim was inaccurate; corrected here and in ADR-0257.
Exporter (placeholder) ai/scripts/gen_mobilesal_placeholder_onnx.py
Registry entry mobilesal_placeholder_v0 in model/tiny/registry.json (smoke=true)
Status Legacy smoke placeholder — superseded for production by saliency_student_v2 (ADR-0444) / saliency_student_v1 (ADR-0286)

The placeholder ONNX is deterministic (no doc_string, fixed producer_version, deterministic protobuf serialisation) so the sha256 stays stable across re-runs of the export script.

For content-dependent saliency, point the extractor at the production default model/tiny/saliency_student_v2.onnx (or model/tiny/saliency_student_v1.onnx). The placeholder is retained to keep the historical ABI / I/O-contract smoke path available.

Input / output contract

The C extractor binds tensors by name, so any future drop-in (real upstream MobileSal export, distilled student, etc.) must declare the exact same names:

inputs:
  input         float32[1, 3, H, W]   ImageNet-normalised RGB, NCHW
outputs:
  saliency_map  float32[1, 1, H, W]   per-pixel saliency in [0, 1]

H and W are dynamic — both the placeholder and the upstream graph match whatever resolution the C side feeds. ImageNet normalisation (mean [0.485, 0.456, 0.406], std [0.229, 0.224, 0.225]) is applied in the C side via the shared vmaf_tensor_from_rgb_imagenet() helper, identical to LPIPS's wiring (see lpips_sq_v1.md).

Usage — CLI

vmaf \
    --reference ref.yuv \
    --distorted dist.yuv \
    --width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
    --feature mobilesal=model_path=model/tiny/saliency_student_v2.onnx \
    --output score.json

The --feature argument takes name=option=value (colon-separated for more options); there is no --feature_params option. The output JSON gains a per-frame saliency_mean column alongside any other features requested in the same run. Combine with lpips and vmaf for the full saliency-quality picture:

vmaf --reference ref.yuv --distorted dist.yuv \
    --width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
    --feature vmaf \
    --feature lpips=model_path=model/tiny/lpips_sq.onnx \
    --feature mobilesal=model_path=model/tiny/saliency_student_v2.onnx \
    --output combined.json

Equivalently, set the model path via env var:

VMAF_MOBILESAL_MODEL_PATH=model/tiny/saliency_student_v2.onnx \
    vmaf --reference ref.yuv --distorted dist.yuv \
        --width 1920 --height 1080 --pixel_format 420 --bitdepth 8 \
        --feature mobilesal --output score.json

Usage — C API

#include <libvmaf/libvmaf.h>

VmafFeatureDictionary *opts = NULL;
vmaf_feature_dictionary_set(&opts, "model_path", "model/tiny/saliency_student_v2.onnx");
int err = vmaf_use_feature(ctx, "mobilesal", opts);
/* ... vmaf_score_pooled(ctx, ..., "saliency_mean", ...) for the per-frame mean */

Equivalent to setting VMAF_MOBILESAL_MODEL_PATH before vmaf_use_feature(ctx, "mobilesal", NULL).

Known limitations

Limit Behaviour Workaround
Bit depth 8-bit YUV only; other depths are rejected at init() with -ENOTSUP (see the message below) Drop --feature mobilesal from HDR / 10-bit / 12-bit runs, or use --bitdepth 8 when the source is 8-bit content in a 10-bit container
Pixel format YUV420P, YUV422P, YUV444P accepted; YUV400P (luma-only) rejected at init() because the model needs three RGB channels Convert to a chroma-carrying format
Colour space BT.709 limited-range Y'CbCr to RGB on the C side, matching feature_lpips.c; BT.2020 / full range is approximate (deliberate trade-off, see the feature_mobilesal.c comment) None
Resolution Bounded by the selected ONNX graph's dynamic shape; the placeholder has no useful quality floor, the fork-trained student was trained on 256x256 crops. The student checkpoints need sides that are multiples of 8; the extractor pads other sizes (576x324 runs as 576x328, the last column and row repeated) and averages the map over the frame's own area (ADR-1540) None for the extractor; pad yourself when running the ONNX graph directly
Execution provider The extractor opens its session with the default auto device: CUDA, OpenVINO GPU, ROCm, CoreML, then CPU, whichever the linked ONNX Runtime provides; independent of -Denable_cuda and of --tiny-device None needed
Score interpretation With the placeholder, saliency_mean is about 0.5 whatever the input; the placeholder only locks down the pipeline. With a student checkpoint the score is content-dependent Use a student checkpoint

With --bitdepth 10 or 12 together with --feature mobilesal the run aborts before scoring. The saliency model requires 8-bit ImageNet-normalised RGB; wider depths would require retraining and are not planned (ADR-0613 §P1-3). The message reads:

mobilesal: bpc=10 is not supported (8-bit only). The mobilesal extractor
requires 8-bit YUV input because the saliency model was trained on 8-bit
ImageNet-RGB. Use --bitdepth 8 or omit --feature mobilesal for HDR /
10-bit / 12-bit content.

Training and evaluation

The placeholder is not trained: it is a generated Conv + Sigmoid with fixed weights, so no training data, quality evaluation or correlation figure exists for it. Trained, evaluated saliency weights are on the saliency_student_v1 and saliency_student_v2 cards.

How the placeholder is regenerated

python ai/scripts/gen_mobilesal_placeholder_onnx.py          # rewrite model/tiny/mobilesal.onnx
python ai/scripts/gen_mobilesal_placeholder_onnx.py --check   # exit 1 if the file differs
# wrote model/tiny/mobilesal.onnx (330 bytes, sha256 f12263...)

The generator writes the ONNX file only; the sidecar JSON and the registry entry are committed files. The output is byte-identical to the shipped file, and --check fails when it is not. The sha256 in registry.json is verified before CreateSession.

  • lpips_sq_v1.md — sister full-reference DNN extractor; shares the YUV → ImageNet-RGB plumbing.
  • ../roadmap.md §2.3 — Wave 1 MobileSal scope.
  • saliency_student_v2.md — production default saliency weights for this extractor (ADR-0444).
  • saliency_student_v1.md — initial fork-trained saliency student baseline (ADR-0286).
  • ADR-0218 — design notes (smoke-only placeholder, scoring-vs-encoder split, scalar-vs-map output).
  • ADR-0257 — first blocker (T6-2a-followup real-weights swap deferred): upstream MobileSal license, distribution and RGB-D mismatch.
  • Research-0053 — upstream survey, licence analysis, and alternatives walk.
  • ADR-0265 — second blocker: U-2-Net u2netp distribution + op-allowlist mismatch.
  • Research-0055 — companion survey for ADR-0265.
  • ADR-0286 — fork-trained production saliency-student path.
  • ADR-0042 — tiny-AI doc-substance rule this page satisfies.