Skip to content

[WIP] gpu-kubelet-plugin: report device health to the kubelet (KEP-4680) - #1525

Draft
harche wants to merge 1 commit into
kubernetes-sigs:mainfrom
harche:kep-4680-device-health-status
Draft

harche wants to merge 1 commit into
kubernetes-sigs:mainfrom
harche:kep-4680-device-health-status

Conversation

@harche

@harche harche commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

What type of PR is this?

/kind feature

What this PR does / why we need it:

Implements KEP-4680 (DRA device health status) for the GPU kubelet plugin, following #1243 (comment).

flowchart LR
  NVML["NVML events"] --> MON["NVML health monitor"]
  MON --> T["Device taints"]
  T --> RS["ResourceSlice (scheduler)"]
  T -->|"on change, every 10s"| W["WatchHealthStatus"]
  W -->|"DRAResourceHealth gRPC"| K["kubelet"]
  K --> P["Pod status: allocatedResourcesStatus"]
Loading

Which issue(s) this PR is related to:

Supersedes #1243.

Special notes for your reviewer:

Two additional changes were needed to test this. Should they stay in this PR or go into separate ones?

  1. GCP nvkind harness on Kubernetes 1.37: kind v0.33.0 and nvkind main. The current harness can't create 1.36+ clusters.
  2. Mock NVML pin bump: k8s-test-infra to 57ef0165 with GPU_COUNT=4. The current pin doesn't deliver injected XIDs, which the new mock-NVML bats test needs.

AI disclosure: an AI coding assistant helped write this change.

Does this PR introduce a user-facing change?

The GPU kubelet plugin reports allocated GPU health in the pod status (KEP-4680) when the NVMLDeviceHealthCheck feature gate is enabled.

Additional documentation (design docs, usage docs, etc.):

site/content/docs/guides/gpu-health-checking.md

@kubernetes-prow

Copy link
Copy Markdown
Contributor

Skipping CI for Draft Pull Request.
If you want CI signal for your change, please convert it to an actual PR.
You can still manually trigger a test run with /test all

@kubernetes-prow kubernetes-prow Bot added do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. release-note Denotes a PR that will be considered when it comes time to generate release notes. kind/feature Categorizes issue or PR as related to a new feature. labels Oct 5, 2026
@netlify

netlify Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for dra-driver-nvidia-gpu canceled.

Name Link
🔨 Latest commit 4e9f612
🔍 Latest deploy log https://app.netlify.com/projects/dra-driver-nvidia-gpu/deploys/6ac3e063358b9f00082592a1

@kubernetes-prow

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by: harche
Once this PR has been reviewed and has the lgtm label, please assign varunrsekar for approval. For more information see the Code Review Process.

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@kubernetes-prow kubernetes-prow Bot added cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files. labels Oct 5, 2026
@harche
harche force-pushed the kep-4680-device-health-status branch from 46699ac to 3ec0e31 Compare October 5, 2026 17:26
@harche
harche force-pushed the kep-4680-device-health-status branch 2 times, most recently from 932e273 to a3c10b3 Compare October 5, 2026 17:35
Implement WatchHealthStatus so the kubelet can surface the health of
allocated GPUs in pod.status.containerStatuses[].allocatedResourcesStatus.

Health is derived from the device taints the NVML health monitor already
publishes in the ResourceSlice, so the pod status and the scheduler see the
same classification: fatal XIDs and GPU loss are Unhealthy, unmonitored
devices are Unknown, non-fatal XIDs stay Healthy with a message. Reports are
sent on subscribe, on every taint change and every 10 seconds, and never
wait behind a Prepare or Unprepare holding the device state lock.

The DRAResourceHealth service is only advertised when the NVML health
monitor runs (NVMLDeviceHealthCheck).

Move the GCP nvkind e2e harness to Kubernetes 1.37, where ResourceHealthStatus
is on by default: use kind v0.33.0 (kubeadm v1beta4) and rebuild nvkind at
main (NVIDIA/nvkind#76, nvidia-ctk --config-source=file) against it. Also keep
macOS tar from adding AppleDouble files that Helm fails to parse as CRDs.

Bump the mock NVML (NVIDIA/k8s-test-infra) to 57ef0165, which delivers
injected XID events to NVML health monitors (NVIDIA/k8s-test-infra#735), and
emulate 4 GPUs to match its 4-GPU gb200 profile; GPUs beyond the profile get
a random UUID per process, so nvidia-smi cannot find them.
@harche
harche force-pushed the kep-4680-device-health-status branch from a3c10b3 to 4e9f612 Compare October 5, 2026 17:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cncf-cla: yes Indicates the PR's author has signed the CNCF CLA. do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. kind/feature Categorizes issue or PR as related to a new feature. release-note Denotes a PR that will be considered when it comes time to generate release notes. size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files.

Projects

Status: Backlog

Development

Successfully merging this pull request may close these issues.

1 participant