Tiny-AI int8 quantisation¶
Pick a quantisation mode, produce a .int8.onnx next to the fp32 model, and check it against an accuracy budget. The fork supports three post-training quantisation (PTQ) modes plus quantisation-aware training (QAT). Each model records its quant decision in model/tiny/registry.json, with a PLCC budget the CI harness enforces against the fp32 baseline.
Policy origin is ADR-0129, audited and scaffolded in ADR-0173. The runtime .int8.onnx redirect landed in ADR-0174. What the loader does today is under How the loader treats int8.
Note
Every shipped int8 file is dynamic PTQ in QOperator format (learned_filter_v1, nr_metric_v1, vmaf_tiny_v3, vmaf_tiny_v4). Static PTQ and QAT are supported by the scripts and the loader, but no shipped registry row uses quant_mode: "static" or "qat".
Choose a mode¶
| Mode | Accuracy | Cost to produce | Best for |
|---|---|---|---|
fp32 | reference | none | new models, debug builds |
dynamic | small hit (~0.5%) | one CLI call | models without a calibration set; deployment box differs from training box |
static | small hit (~0.2%) | one calibration pass | models you own and can pin a calibration set for |
qat | reference (within ~0.05%) | extra training phase, ~1.5x fp32 train time | models where static drops accuracy past the per-model budget |
Pick the cheapest mode that stays inside the quant_accuracy_budget_plcc budget.
Registry fields¶
| Field | Type | Default | Required when |
|---|---|---|---|
quant_mode | fp32 / dynamic / static / qat | fp32 | always present (default fp32) |
quant_calibration_set | path relative to the repo root | absent | quant_mode == "static" |
quant_accuracy_budget_plcc | number in [0, 1] | 0.01 | always (the CI gate honours per-entry values) |
fp32 keeps the loader on <basename>.onnx. The other three modes redirect the loader to a sibling <basename>.int8.onnx produced by the scripts below. The fp32 file stays on disk as the regression baseline. The full schema is in model-registry.md.
Produce int8 artefacts¶
Dynamic PTQ¶
No calibration data is needed. The script wraps onnxruntime.quantization.quantize_dynamic.
python ai/scripts/ptq_dynamic.py model/tiny/nr_metric_v1.onnx \
--report-out runs/nr_metric_v1_dynamic_ptq.json
# -> model/tiny/nr_metric_v1.int8.onnx
--report-out writes a JSON report with the fp32 and int8 byte sizes, the per-channel setting, the output path and run_provenance.
Static PTQ¶
- Build a calibration
.npz: one entry per ONNX input name, each a stack of[N, ...]representative samples. No in-tree script writes it; hand-craft it from a parquet feature cache or decoded frames. For the feature-vector FR regressors,vmaf-train quantize-int8(see training.md) runs static PTQ calibrated from a parquet feature cache directly, without an.npz. -
Quantise:
-
The output goes to
<input>.int8.onnx. Add the calibration path to the registry'squant_calibration_setfield.
The optional report holds the calibration input names and sample count, the size ratio and run_provenance. No calibration .npz is shipped: calibration sets are not redistributable by default, so each operator builds their own.
vmaf-train quantize-int8¶
A one-command static PTQ (QDQ format) from a parquet feature cache, with a drift gate. It reports the int8-versus-fp32 RMSE on held-out samples and exits 2 when the RMSE exceeds the gate.
vmaf-train quantize-int8 \
--fp32 runs/fr_tiny_v1/fr_tiny_v1.onnx \
--output runs/fr_tiny_v1/fr_tiny_v1.int8.onnx \
--calibration ai/data/nflx_features.parquet \
--json runs/fr_tiny_v1/quantize_int8.json
| Flag | Default | Meaning |
|---|---|---|
--fp32 | required | Input fp32 .onnx. |
--output | required | Output int8 .onnx. |
--calibration | required | Parquet feature cache used for calibration (the schema vmaf-train eval consumes). |
--input-name | features | ONNX input name. |
--n-calibration | 512 | Calibration sample count. |
--batch-size | 32 | Calibration batch size. |
--rmse-gate | 1.0 | Exit 2 when the int8-versus-fp32 RMSE exceeds this. |
--json PATH | unset | Write a JSON report with run_provenance. |
Quantisation-aware training (QAT)¶
Use QAT when static PTQ exceeds the per-model quant_accuracy_budget_plcc, or when the QAT-versus-static delta on real content justifies about 50 % extra training time (Research-0006 section 4). On tiny models (about 10 K parameters and fewer layers) QAT and static PTQ tend to agree inside the 0.002 budget, so choose static PTQ for cost. On larger architectures with wider weight distributions QAT typically wins. The measured delta is recorded per model in its ADR, for example ADR-0208.
python ai/scripts/qat_train.py \
--config ai/configs/learned_filter_v1_qat.yaml \
--output model/tiny/learned_filter_v1.int8.onnx \
--report-out runs/learned_filter_v1_qat.json
Phases¶
ADR-0207 defines four phases:
| Phase | What happens |
|---|---|
| 1. fp32 warm-start | Normal fp32 training. |
| 2. Fake-quant insertion | torch.export captures the trained module and torchao.quantization.pt2e.prepare_qat_pt2e inserts observers under X86InductorQuantizer's default recipe: per-tensor uint8 activations, per-channel symmetric int8 weights on ch_axis=0. |
| 3. QAT fine-tune | Fine-tune at 10x reduced learning rate (default \(\mathrm{fp32\_lr} / 10\)). |
| 4. ONNX export | Copy the QAT-conditioned weights into a fresh fp32 module, export that graph, then run onnxruntime.quantization.quantize_static with a calibration set drawn from the QAT training distribution. |
The output is a QDQ-format .int8.onnx, structurally identical to the static-PTQ artefact. The QAT effect is preserved entirely through weight pre-conditioning.
CLI knobs¶
| Flag | Default | Meaning |
|---|---|---|
--epochs-fp32 | 20 | Phase 1 epochs. |
--epochs-qat | 10 | Phase 3 epochs. |
--lr-qat | fp32 lr / 10 | Phase 3 learning rate. |
--n-calibration | 64 | Calibration samples for phase 4. |
--smoke | off | Skip both training phases (CI / dev round trip). |
--report-out | unset | JSON with fp32/int8 outputs, parameter count, phase settings and run_provenance. |
The YAML config mirrors the vmaf-train fit shape plus a qat: block; a complete example is ai/configs/learned_filter_v1_qat.yaml. ai.train.qat.run_qat(...) exposes the same pipeline for direct Python use, as tests do.
Training data¶
The config's cache: field points at the training corpus. The rank of qat.input_shape decides which reader parses it, the same shape the trace uses, so loader and trace cannot disagree:
qat.input_shape rank | Loader | Cache format |
|---|---|---|
2 (for example [1, 6]) | vmaf_train.datamodule.VmafTrainDataModule | .parquet (canonical-6 feature columns plus mos) or .npz (features, scores) |
4 (for example [1, 1, 32, 32]) | built-in NCHW image loader | .npz (below) |
| anything else | none; the run downgrades to --smoke with a message on stderr | none |
Rank 4 is also the default when qat.input_shape is absent, matching the [1, 1, 32, 32] fallback the exporter traces with.
A rank-4 .npz carries the input batch under x (aliases images, degraded, input) with shape (N, C, H, W), and the target under y (aliases targets, clean, reference, output) with the same leading dimension. Both are cast to float32. The archive is validated when the loader is built: a wrong extension, an unknown array name, a non-4D input, a length mismatch or an empty cache all exit before the fp32 warm-start burns an epoch. Batch size comes from the config's batch_size (default 32).
import numpy as np
np.savez(
"ai/data/learned_filter_patches.npz",
x=degraded_patches, # (N, 1, 32, 32) float32
y=clean_patches, # (N, 1, 32, 32) float32
)
How the loader treats int8¶
Both vmaf_dnn_session_open in core/src/dnn/dnn_api.c and vmaf_use_tiny_model in core/src/dnn/dnn_attach_api.c use the same redirect logic to choose the file:
| Step | Check | On failure |
|---|---|---|
| 1 | Validate the caller-supplied fp32 path: size cap and op allowlist. | Error. |
| 2 | Load the sidecar. If the caller path already ends in .int8.onnx, or quant_mode: "fp32", the given file is the model. | Resolution stops here on success. |
| 3 | Strip a trailing .onnx, append .int8.onnx, run the same size and allowlist validation on that sibling. | A path overflowing the 4096-byte buffer returns -ENAMETOOLONG. |
| 4 | If the int8 file validates, load it instead of the fp32 file. | Go to step 5. |
| 5 | The int8 file is missing, over the size cap, or has a non-allowlisted op. | Log a VMAF_LOG_LEVEL_DEBUG line and load the fp32 baseline. The session still reports the sidecar's quant_mode; only the weights are fp32. |
| 6 | The int8 file passed step 3 but vmaf_ort_open() fails on it. | Retry the fp32 baseline once. A failure of that retry reaches WARNING. |
The redirect keys off quant_mode != fp32 alone. It does not tell dynamic from static or qat, and neither registry nor sidecar records a wire format. The allowlist scan in step 3 alone decides whether an int8 graph is acceptable.
Step 6 is real. An ONNX Runtime build without a kernel for a quantised op fails session creation with -EIO and Could not find an implementation for ConvInteger(10). model/tiny/nr_metric_v1.int8.onnx (dynamic PTQ, ConvInteger) hits this on such a build. The allowlist scan cannot see it, because it checks op names, not whether the local runtime has a kernel. The int8 attempt's error is logged at DEBUG, like step 5. Both loaders share step 6 through vmaf_ort_open_with_fallback() in core/src/dnn/ort_backend.c.
Warning
A quantised model whose int8 file is absent or rejected still loads and still scores, at fp32 weights and fp32 speed. Nothing fails, and the scores are the fp32 baseline's rather than the int8 model's. The fallback is visible only at debug log level, and the vmaf CLI has no flag for it (the binary runs at VMAF_LOG_LEVEL_INFO). An API caller that sets VmafConfiguration.log_level to VMAF_LOG_LEVEL_DEBUG sees int8 sidecar unavailable (steps 3 to 5) or int8 session open failed (step 6) and can confirm which weights a session loaded.
Integrity of the int8 file¶
The redirect does not verify int8_sha256. The fp32 sidecar records the digest for its quantised sibling, but neither loader parses it. At load time the gates are the 50 MB size cap and the op allowlist. Integrity is enforced elsewhere:
ai/scripts/validate_model_registry.pyandcore/test/dnn/test_registry.shcompareint8_sha256against the file on disk in CI;--tiny-model-verifychecks the Sigstore bundle (ADR-0211, model-registry.md).
A digest check inside the loader would change the ADR-1032 fallback semantics (a mismatch would need its own outcome, distinct from "int8 absent") and needs its own ADR. It is not in the tree.
The onnx_has_scaler contract¶
Feature-vector tiny models (vmaf_tiny_v2 to v4, fr_regressor_v*) ship the StandardScaler in one of two ways, and the sidecar is the only thing that says which:
| Where the scaler lives | Sidecar | Runtime |
|---|---|---|
In the graph: Sub (mean) and Div (std) Constant nodes before the first Gemm | declares "onnx_has_scaler": true | feeds raw canonical-6 values |
| In the runtime | carries input_mean / input_std (or feature_mean / feature_std), omits onnx_has_scaler | core/src/libvmaf.c normalises the vector before inference |
The two must not both apply. If a scaler-baking graph ships without the declaration, the runtime double-scales and the score is meaningless. Measured on the Netflix src01_hrc00/hrc01_576x324 pair with vmaf_tiny_v3.int8.onnx:
| Sidecar | Pooled vmaf_tiny_model | Per-frame PLCC vs fp32 |
|---|---|---|
| without the declaration | 16.020865 | 0.975443 |
| with the declaration | 71.952113 | 0.999876 |
| fp32 baseline | 72.359458 | 1 |
Quantising a model does not change which case applies: ptq_dynamic.py, ptq_static.py and qat_train.py keep the scaler nodes in the graph. An .int8.onnx therefore needs the same declaration its fp32 parent has. Three gates enforce it over every model/tiny/*.int8.onnx:
bash core/test/dnn/test_registry.sh # meson `dnn` suite
python -m pytest python/test/model_registry_schema_test.py -q # python suite
python ai/scripts/validate_model_registry.py # CI registry gate
Each loads the graph with onnx when installed (falling back to a protobuf byte scan) and fails when a graph with both Sub and Div has a companion sidecar that does not declare "onnx_has_scaler": true. The companion sidecar of foo.int8.onnx is foo.int8.json when present, otherwise the fp32 sidecar foo.json, the same order vmaf_dnn_sidecar_load uses.
Wire format: QOperator versus QDQ¶
ONNX encodes an int8 graph in one of two wire formats, and the fork does not treat them interchangeably. quant_mode records how a model was quantised, not the format. The format follows from the producer script:
| Producer | quant_mode | Wire format | Int8 ops emitted |
|---|---|---|---|
ai/scripts/ptq_dynamic.py | dynamic | QOperator | DynamicQuantizeLinear, MatMulInteger, ConvInteger |
ai/scripts/ptq_static.py | static | QDQ (pinned) | QuantizeLinear / DequantizeLinear wrapping stock Conv / Gemm / MatMul |
ai/src/vmaf_train/quantize.py (vmaf-train quantize-int8) | static | QDQ (pinned explicitly) | as above, restricted to Gemm / MatMul / Conv |
ai/scripts/qat_train.py | qat | QDQ | as above; the export phase runs quantize_static |
What the fork ships¶
QOperator, for every model shipped today. All four .int8.onnx files under model/tiny/ come from ptq_dynamic.py, and ONNX Runtime's quantize_dynamic takes no quant_format argument, so dynamic PTQ is QOperator-only. No shipped model contains a QuantizeLinear or DequantizeLinear node: QDQ is a supported input, not a shipped output.
Node census of the shipped int8 files
What the loader accepts¶
QOperator dynamic and QDQ load. QOperator static is rejected.
The gate is core/src/dnn/op_allowlist.c. vmaf_dnn_scan_onnx walks every ONNX file node by node, recursively into Loop and If subgraphs. A node whose op_type is not on the list makes vmaf_dnn_validate_onnx return -EPERM. The five quantisation entries are:
| Op | Format | Role |
|---|---|---|
QuantizeLinear | QDQ | fp32 to int8 with a calibrated scale and zero-point |
DequantizeLinear | QDQ | int8 back to fp32 on leaving a quantised region |
DynamicQuantizeLinear | QOperator dynamic | per-tensor scale and zero-point computed at run time |
MatMulInteger | QOperator dynamic | integer matrix multiply |
ConvInteger | QOperator dynamic | integer convolution |
QDQ loads because it leaves the arithmetic on stock Conv, Gemm and MatMul nodes, which the allowlist already carries for fp32 models. The QDQ pair adds only the two scale-carrying ops.
QOperator static does not load. It folds the arithmetic into fused QLinear* ops (QLinearConv, QLinearMatMul, QGemm, QLinearAdd and others) and the allowlist has none of them. This is why ai/src/vmaf_train/quantize.py pins quant_format=QuantFormat.QDQ instead of relying on a default. Admitting a QLinear* op would widen the model attack surface and is an allowlist change that needs a security review, not a documentation change.
Accuracy gates¶
Two gates guard int8 quality. The first runs in CI, the second is the strict clip-level check for release quality.
| Gate | Script | Input | Threshold | Exit codes |
|---|---|---|---|---|
ai-quant-accuracy (CI) | ai/scripts/measure_quant_drop.py | 16 deterministic synthetic samples (seed 0) | PLCC drop at most the per-model quant_accuracy_budget_plcc | 0 pass, 1 over budget or missing file, 2 bad invocation |
| Clip-level parity | ai/scripts/validate_quant_parity.py | real feature clips (default testdata/scores_cpu_576.json; .parquet and .npz supported) | mean absolute delta at most 0.10 VMAF (--max-mean-delta), maximum single-frame delta at most 0.50 (--max-single-delta), PLCC at least 0.990 (--min-plcc) | 0 all pass, 1 any threshold breached, 2 invocation or file errors |
CI gate: ai-quant-accuracy¶
The job runs in the Tiny AI job of tests-and-quality-gates.yml (wired by ADR-0174). It calls measure_quant_drop.py --all, which walks the registry, runs each non-fp32 model through fp32 and int8 ORT sessions on 16 deterministic synthetic samples, and asserts that the aggregate Pearson correlation drop stays below the per-model budget. A budget violation fails the PR. Run it locally:
--out-json keeps the per-model gate rows and the run_provenance block. Use it as model-card evidence, or to compare a refreshed int8 sidecar with a previous CI gate.
Gate a model that is not in the registry¶
--all and the positional form resolve through model/tiny/registry.json: the model must live under model/tiny/ and the budget comes from its entry. A model that is not committed yet (PTQ or QAT scratch output, a CI artefact, a release candidate) has neither. The --fp32 and --int8 overrides measure an explicit pair and touch no registry:
python ai/scripts/measure_quant_drop.py \
--fp32 /tmp/train_out/mlp_small_final.onnx \
--int8 /tmp/train_out/mlp_small_final.ptq_static.int8.onnx \
--budget 0.002 \
--id mlp_small_static_ptq \
--out-json /tmp/quant_drop.json
[PASS] mlp_small_static_ptq mode=override PLCC=0.999536 drop=0.000464 budget=0.0020 worst_abs=0.0010
| Flag | Meaning |
|---|---|
--fp32 PATH | fp32 ONNX to measure. Required together with --int8; any path. |
--int8 PATH | int8 ONNX to measure against it. |
--budget FLOAT | PLCC-drop budget for the pair (default 0.01, the registry-wide default). Research-2029 section 6 recommends 0.002 for static PTQ and 0.001 for QAT. |
--id NAME | Label in the console line and the report. Defaults to the fp32 filename without .onnx. |
The overrides are mutually exclusive with --all and with the positional argument: passing both exits 2. The report keeps the registry-run shape, with "quant_mode": "override" on the single row.
Clip-level gate: validate_quant_parity.py¶
Synthetic random tensors test operator fidelity but are less sensitive than real natural-video features. validate_quant_parity.py evaluates models on real feature clips against the strict Research-2029 section 6 thresholds in the table above.
# Gate all shipped tabular FR models against the default thresholds:
python ai/scripts/validate_quant_parity.py --all \
--out-json runs/quant_parity_gate.json
# Gate a specific model or a direct fp32/int8 pair:
python ai/scripts/validate_quant_parity.py --model vmaf_tiny_v3
python ai/scripts/validate_quant_parity.py \
--fp32 /tmp/model.onnx \
--int8 /tmp/model.int8.onnx \
--features ai/testdata/bisect/features.parquet \
--max-mean-delta 0.10 \
--max-single-delta 0.50 \
--min-plcc 0.990
Note
The shipped vmaf_tiny_v3 and vmaf_tiny_v4 are un-retrained dynamic PTQ models. They reach high linear correlation (\(\text{PLCC} \ge 0.994\)), but dynamic activation quantisation adds a modest absolute score shift (\(\text{mean } |\Delta| \approx 0.31\text{--}0.56\), \(\text{max } |\Delta| \approx 0.54\text{--}0.97\)). Running validate_quant_parity.py with the default thresholds therefore fails closed, as designed. Reaching mean delta \(\le 0.10\) and max delta \(\le 0.50\) needs full QAT retraining against the vmaf_v1.0.16_3d0h teacher on the 152k-clip training corpus (Epic #1246, RC9).
Currently quantised models¶
| Model id | Mode | Size shrink | Measured drop | Budget |
|---|---|---|---|---|
learned_filter_v1 | dynamic | 2.4x (80 KB to 33 KB) | 0.000117 (PLCC 0.999883) | 0.01 |
nr_metric_v1 | dynamic | 2.0x (119 KB to 58 KB) | 0.007674 (PLCC 0.992326) | 0.01 |
vmaf_tiny_v3 | dynamic | 0.95x (4 496 B to 4 267 B) | 0.000120 (PLCC 0.999880) | 0.01 |
vmaf_tiny_v4 | dynamic | 1.8x (14 046 B to 7 769 B) | 0.000145 (PLCC 0.999855) | 0.01 |
vmaf_tiny_v3 and vmaf_tiny_v4 joined the dynamic-PTQ family in ADR-0275. Their model cards carry the reproduction commands and measured PLCC drops: vmaf_tiny_v3 and vmaf_tiny_v4.
Propose a model for quantisation¶
- Run
ai/scripts/ptq_<mode>.pyto produce the int8 file. - Compute fp32 versus int8 PLCC on the soak fixture.
- In the PR description, paste the PLCC numbers and the fp32 / int8 inference time ratio on at least one CPU.
- Update
model/tiny/registry.json: flipquant_modeto the chosen mode, setquant_accuracy_budget_plcc(default 0.01, one PLCC point), and addquant_calibration_setifstatic. - Land the int8 ONNX next to the fp32 file.
The reviewer compares the measured drop against the budget. If a static run misses the budget, escalate to QAT in a follow-up PR; do not relax the budget.
Caveats¶
ptq_static.pypinsquant_format=QuantFormat.QDQ. Likeai/src/vmaf_train/quantize.py, it pins QDQ so emitted static graphs contain only allowlisted ops. A QOperator-static graph (QLinearConv,QLinearMatMul,QGemm, ...) would be rejected bycore/src/dnn/op_allowlist.cand libvmaf would silently fall back to fp32.test_ptq_static_full_roundtripinai/tests/test_ptq_scripts.pytests this contract.- Calibration sets are not redistributable by default. Operators build their own from a parquet feature cache.
- int8 can be slower than fp32. VNNI / DLBoost speedup applies only to Intel CPUs from Cascade Lake on; ARMv8.2 and later have int8 dot-product. Without either, the int8 path can run slower than fp32. The overhead is the QOperator-dynamic requantise chain, not QDQ (no shipped model contains a
QuantizeLinearnode). EveryMatMulInteger/ConvIntegeris preceded by aDynamicQuantizeLinearthat recomputes scale and zero-point on every inference, and followed by aCast/Mul/Addchain converting the int32 accumulator to fp32. Without an integer dot-product instruction these are pure overhead. The loader is bit-depth-agnostic and still picks the int8 model when the registry says so; measuring runtime performance is the operator's job.
History¶
- 2026-09-05. The
onnx_has_scalerdouble-scaling defect was fixed inmodel/tiny/vmaf_tiny_v3.int8.json, which had shipped without the declaration (T-TINY-V3-INT8-SIDECAR-MISSING-ONNX-HAS-SCALER-2026-09-04). The three gates above were added to prevent a repeat. - ADR-1293. QAT moved off
torch.ao.quantization.quantize_fx.prepare_qat_fx, which PyTorch deprecated wholesale (the repository treats itsDeprecationWarningas an error), totorchaoPT2E (ADR-1293). Weight quantisation is unchanged. The activation range widened from the old reduce-range [0, 127] to the full [0, 255], which ORTquantize_statichas always used on the other side of the handoff. Phase 4 also moved to thetorch.export-based ONNX exporter: the export target is a plain fp32 module, so the quantisation buffers that forced the legacy TorchScript path are gone. - ADR-1032, fp32 fallback. ADR-0174 section 2 originally specified that a missing int8 file returns a negative error ("no silent fp32 fallback, that would mask deployment misconfigurations"). ADR-1032 Fix 3 replaced the hard error with the fallback on a "better degraded than dead" rationale, and the code has matched ADR-1032 since. ADR-0174 is Accepted and therefore frozen, so it still reads the old way. This page is authoritative for the runtime behaviour.
- QAT data dispatch. Before the rank-based loader dispatch, every config went to
VmafTrainDataModule, which only materialises rank-2 tabular rows. 2D CNN configs such aslearned_filter_v1_qat.yamlcould not train for real. They only appeared to work because their parquet cache is uncommitted, so the missing-cache branch silently downgraded the run to smoke mode (Research-2029 section 5 gap 4). nr_metric_v1export (T5-3d). The original ONNX export tripped ORT shape inference inquantize_dynamicwithInferred shape and existing shape differ in dimension 0: (128) vs (1).torch.onnx.exporthad emitted every initialiser intograph.value_infowith static-shape annotations that did not survive the dynamic batch axis. The exporter (ai/src/vmaf_train/models/exports.py) andai/scripts/ptq_dynamic.pynow strip those duplicates, the same workaround introduced forvmaf_tiny_v1*.onnxin PR #174 (T5-3e).