vmaftune.predictor: Python API reference¶
Use vmaftune.predictor to estimate the VMAF of a shot at a given CRF from cheap probe-encode signals, and to invert that estimate into the CRF that meets a target VMAF, without running VMAF on the final encode. It is the predict half of the predict-then-verify loop; vmaftune.predictor_validate is the verify half. The source is tools/vmaf-tune/src/vmaftune/predictor.py.
For the model, training and the production-flip gate see the predictor guide. Interval prediction is described in conformal VQA.
Note
The navigation entry for this page reads vmaftune.predict; the module is vmaftune.predictor.
Quick start¶
This program runs without onnxruntime or a model file: Predictor() uses the built-in analytical curve.
from vmaftune.predictor import Predictor, ShotFeatures
features = ShotFeatures(
probe_bitrate_kbps=3200.0,
probe_i_frame_avg_bytes=48000.0,
probe_p_frame_avg_bytes=9000.0,
probe_b_frame_avg_bytes=4000.0,
shot_length_frames=240,
fps=24.0,
width=1920,
height=1080,
)
predictor = Predictor()
vmaf = predictor.predict_vmaf(features, 28, "libx264") # CRF 28
crf = predictor.pick_crf(features, 93.0, "libx264") # largest CRF with VMAF >= 93
keyint, min_keyint = predictor.pick_keyint(features, 24.0)
With these inputs predict_vmaf returns about 93.96, pick_crf returns 28 and pick_keyint returns (48, 12).
ShotFeatures¶
A frozen dataclass of cheap signals for one shot. Each is read from the probe encode's output or from a single centre frame; none needs VMAF or the final encode. All values are non-negative.
| Field | Type | Default | Meaning |
|---|---|---|---|
probe_bitrate_kbps | float | required | Average bitrate of the probe encode. |
probe_i_frame_avg_bytes | float | required | Mean I-frame size. |
probe_p_frame_avg_bytes | float | required | Mean P-frame size. |
probe_b_frame_avg_bytes | float | required | Mean B-frame size; 0 when the codec has none. |
saliency_mean | float | 0.0 | Mean saliency in [0, 1] (saliency student model). |
saliency_var | float | 0.0 | Saliency variance across the centre frame. |
frame_diff_mean | float | 0.0 | Mean absolute frame difference (motion energy). |
y_avg | float | 0.0 | Mean luma over the shot. |
y_var | float | 0.0 | Luma variance. |
shot_length_frames | int | 0 | Shot length in frames. |
fps | float | 0.0 | Frame rate. |
width, height | int | 0 | Frame size. |
resolution_class(height) maps a height to the ladder vocabulary: "sd" (up to 480), "hd_ready" (up to 720), "hd" (up to 1080), "uhd" (up to 2160) and "uhd8k" (everything above 2160).
Predictor¶
A dataclass. There is no encoder or preset argument: the codec is chosen per call.
| Argument | Effect |
|---|---|
model_path=None | Analytical fallback only. |
model_path=<predictor_<codec>.onnx> | Loads the learned per-codec model on the CPU execution provider. |
coefficients | Per-codec (a, b, c, d, crf_ref) for the analytical curve; defaults cover 14 codecs (libx264, libx265, libsvtav1, libaom-av1, libvvenc and the NVENC, AMF and QSV H.264, HEVC and AV1 encoders). |
Behaviour to know:
- No silent substitution when onnxruntime is missing. If onnxruntime cannot be imported, the predictor uses the analytical curve. If it can be imported and
model_pathdoes not exist, construction raisesFileNotFoundError; an invalid graph propagates its own error. - Stub models.
is_stub(read-only) isTruewhen the model file name containsstub, when its card (<name>_card.md) sayssynthetic-stub, or, unless the card saysreal-N=, when the name contains a codec that ships only a synthetic model (libx264, libx265, libsvtav1, libaom-av1, libvvenc, h264_amf, hevc_amf, av1_amf). Loading a stub emits aUserWarning: such models are not authoritative for production CRF picks. - Analytical curve. \(\mathrm{vmaf} = a - b\,\delta - c\,\delta^2 + d \log_{10}(\mathrm{bitrate\_kbps})\) with
delta = crf - crf_refand the codec's constants, bitrate floored at 1 kbps and the result clamped to[0, 100]. An unknown codec uses the libx264 constants. The constants are seed values for tests, not trained results.
Methods¶
| Method | Returns | Does |
|---|---|---|
predict_vmaf(features, target_quality, codec) | float in [0, 100] | Predicted VMAF at CRF target_quality (an int) for codec. Uses the ONNX model when loaded, else the analytical curve. |
predict_mos(features, codec, *, target_quality=None) | float in [1, 5] | Predicted MOS. Uses model/konvid_mos_head_v1.onnx when it ships and onnxruntime imports; otherwise (predict_vmaf - 30) / 14, clamped, an approximation. target_quality=None uses the codec's default CRF. |
pick_crf(features, target_vmaf, codec) | int | Binary search over the codec adapter's quality_range for the largest CRF whose predicted VMAF is at least target_vmaf. Returns the adapter's default CRF if none qualifies. |
pick_keyint(features, fps) | (keyint, min_keyint) | GOP heuristic, below. |
predict_vmaf_with_uncertainty(features, target_quality, codec, *, calibration=None, alpha=None) | (point, low, high) | Point estimate with a conformal interval. calibration is a SplitConformalCalibration or CVPlusConformalCalibration from vmaftune.conformal; None returns (point, point, point). alpha overrides the calibration's miscoverage level. |
The interval covers the true value with probability at least 1 - alpha under exchangeability (split conformal, Lei et al. 2018; CV+, Barber et al. 2021).
GOP heuristic¶
pick_keyint(features, fps) returns (keyint, min_keyint) from the probe bitrate (kbps) and the shot length:
| Condition | keyint | min_keyint |
|---|---|---|
| Shot of at least 4 s and probe bitrate below 1500 kbps | \(4 \cdot \mathrm{fps}\) | fps |
| Probe bitrate above 8000 kbps | fps | \(\max(\mathrm{fps} / 2, 1)\) |
| Otherwise | \(2 \cdot \mathrm{fps}\) | \(\max(\mathrm{fps} / 2, 1)\) |
fps is rounded to an integer first (at least 1). The thresholds are constants chosen for 1080p natural content. The learned model picks the CRF on top of these bands.
Module functions¶
| Function | Does |
|---|---|
pick_crf(predictor, features, target_vmaf, codec) | Function form of Predictor.pick_crf. |
pick_keyint(features, fps) | Function form of Predictor.pick_keyint. |
make_predictor_predicate(predictor, feature_extractor) | Adapts a predictor to the PredicateFn seam of vmaftune.per_shot.tune_per_shot: (shot, target_vmaf, encoder) -> (crf, predicted_vmaf). feature_extractor(shot, encoder) returns a ShotFeatures; the real one lives in vmaftune.predictor_features, tests inject a stub. |
from vmaftune.predictor import Predictor, make_predictor_predicate
predicate = make_predictor_predicate(Predictor(), my_feature_extractor)
crf, predicted_vmaf = predicate(shot, 93.0, "libx264")
Uncertainty thresholds¶
vmaftune.uncertainty.ConfidenceThresholds turns an interval width into a decision: tight_interval_max_width (default 2.0 VMAF) and wide_interval_min_width (default 5.0 VMAF). See the ladder API for how the ladder uses them and conformal VQA for the calibration.