ADR-1547: The Helm chart derives the GPU resource name from the vendor and the Intel kernel driver, with an explicit override¶
- Status: Accepted
- Date: 2026-10-04
- Deciders: maintainer, agent
- Tags: helm, k8s, gpu, intel, fork-local
Context¶
The chart requested gpu.intel.com/i915 for every Intel GPU, written out in two helpers of deploy/helm/vmafx/templates/_helpers.tpl (docs audit of 2026-10-03, defect 36). The Intel device plugin names the resource after the GPU's kernel driver: gpu.intel.com/xe for GPUs on xe (Arc B-series and newer, or an older GPU bound to xe). On such a node a VMAFX pod stayed Pending. The maintainer's home cluster advertises gpu.intel.com/xe on its Arc B580 and Pro B60 nodes and gpu.intel.com/i915 on its Arc A380 node.
Decision¶
gpu.intelDriver(i915|xe, defaulti915) selects the Intel resource:gpu.intel.com/<intelDriver>.gpu.resourceNamerequests any extended resource verbatim (validated asdomain/name), overriding the vendor default andintelDriver; it is refused withgpu.vendor: cpu. The backend still followsgpu.vendor.- One helper,
vmafx.gpuResourceName, resolves the name for every workload;vmafx.gpuResourceadds thegpu.enabledgate; the duplicatevmafx.gpuResourceKeyis removed. - The default stays
i915, so a values file that rendered before renders the same manifest.
Alternatives considered¶
| Option | Pros | Cons | Why not chosen |
|---|---|---|---|
Default intelDriver to xe | Matches the newest GPUs | Changes the request of every existing Intel deployment without notice | Default kept; xe documented |
| Request both resources | No choice for the operator | A pod requests what one node never has both of; it would never schedule | Not chosen |
Only gpu.resourceName | One value | The common Intel case needs the plugin's exact name typed by hand, unvalidated | Driver enum for the common case, override for the rest |
| Detect the driver at install time | No configuration | Helm renders without cluster access in CI and GitOps; lookup returns nothing there | Not chosen |
Consequences¶
- Positive: Intel GPUs on
xeschedule withgpu.intelDriver: xe; sharing and MIG resources work throughgpu.resourceName. - Negative: one more value to set for
xenodes. - Neutral / follow-ups:
scripts/ci/tests/test_helm_node_contract.pyrenders the server and node Deployments for every vendor,xe, an override, and the refusals.
References¶
- Docs audit defect list of 2026-10-03, defect 36.
- The maintainer's home-cluster configuration (a private repository, read only): its Intel device plugin advertises
gpu.intel.com/xefor the B580 and B60 nodes. - Popup answer of 2026-10-04, "Defects": "Track all, fix now (Recommended)".