Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 6 additions & 1 deletion hack/ci/gcp-nvkind/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,12 @@ the full env. `TESTINFRA_DIR` points `e2e-test.sh` at the shared lib under
- `registry.k8s.io/dra-driver-nvidia/dra-driver-nvidia-gpu:v0.4.0-dev` is
not published; `install-dra-driver.sh` builds from source and
`kind load docker-image`s the result.
- `kindest/node:v1.34.3` (DRA GA). `v1.34.1` is not published on Docker Hub.
- `kindest/node:v1.37.0`. Device health status (`ResourceHealthStatus`) is
beta and on by default since 1.36, so no extra kubelet feature gate is
needed to exercise it.
- kind `v0.33.0`, and nvkind rebuilt against it: Kubernetes 1.37 requires the
kubeadm v1beta4 config that kind v0.32.0+ generates. See
`lib/setup-nvkind-node.sh`.
- GPU Operator `v26.3.1` in minimal mode: `driver`, `toolkit`, and
`devicePlugin` disabled; `cdi.enabled=true`, `nfd.enabled=true`.
- DLVM ships the NVIDIA driver and toolkit but not Docker/Go/kind/helm/kubectl;
Expand Down
6 changes: 4 additions & 2 deletions hack/ci/gcp-nvkind/e2e-test.sh
Original file line number Diff line number Diff line change
Expand Up @@ -50,7 +50,7 @@ source "${LIB_DIR}/gce.sh"
: "${BOSKOS_HOST:=http://boskos.test-pods.svc.cluster.local}"
: "${BOSKOS_RESOURCE_TYPE:=gpu-project}"
: "${GCE_ZONE:=us-central1-b}"
: "${K8S_VERSION:=v1.34.3}"
: "${K8S_VERSION:=v1.37.0}"
: "${GPU_OPERATOR_VERSION:=v26.3.1}"

mkdir -p "${ARTIFACTS}"
Expand Down Expand Up @@ -88,7 +88,9 @@ gce::wait_for_driver

# 2. Ship repo to VM. Vendor is included (Dockerfile uses -mod=vendor).
WORK=$(mktemp -d)
tar --exclude='.git' --exclude='_output' --exclude='dist' --exclude='.claude' --exclude='site' \
# COPYFILE_DISABLE and the ._* exclude keep macOS tar from adding AppleDouble
# files, which Helm then fails to parse as CRDs on local runs from a Mac.
COPYFILE_DISABLE=1 tar --exclude='._*' --exclude='.git' --exclude='_output' --exclude='dist' --exclude='.claude' --exclude='site' \
-czf "${WORK}/dra-src.tgz" -C "${REPO_ROOT}" .
gce::scp_to "${WORK}/dra-src.tgz" "/tmp/"
gce::ssh 'rm -rf /tmp/dra-src && mkdir -p /tmp/dra-src && tar -xzf /tmp/dra-src.tgz -C /tmp/dra-src'
Expand Down
26 changes: 21 additions & 5 deletions hack/ci/gcp-nvkind/lib/setup-nvkind-node.sh
Original file line number Diff line number Diff line change
Expand Up @@ -105,11 +105,27 @@ if ! command -v helm >/dev/null 2>&1; then
curl -fsSL https://raw.githubusercontent.com/helm/helm/main/scripts/get-helm-3 | bash
fi

# kind (pinned to the version nvkind vendors) and nvkind. nvkind has no
# tags; pin a commit SHA. Bump with a reviewable diff after re-validating.
: "${NVKIND_SHA:=9a3061c75e59ac7e4f29e001f1c5875d46a7cc54}"
go install sigs.k8s.io/kind@v0.31.0
go install "github.com/NVIDIA/nvkind/cmd/nvkind@${NVKIND_SHA}"
# kind and nvkind. nvkind has no tags; pin a commit SHA. Bump with a
# reviewable diff after re-validating.
#
# kind v0.32.0+ is required for Kubernetes 1.37, which dropped the kubeadm
# v1beta3 config that older kind generates. nvkind still vendors kind v0.31.0
# (NVIDIA/nvkind#75 bumps it), so nvkind is rebuilt against KIND_VERSION; the
# pinned nvkind commit includes NVIDIA/nvkind#76, without which the containerd
# in newer kindest/node images fails to restart after nvidia-ctk configures it.
: "${KIND_VERSION:=v0.33.0}"
: "${NVKIND_SHA:=c57050497cffee36f10c8b918d2727c4f26d886e}"
go install "sigs.k8s.io/kind@${KIND_VERSION}"
rm -rf /tmp/nvkind-src
git clone -q https://github.com/NVIDIA/nvkind /tmp/nvkind-src
git -C /tmp/nvkind-src checkout -q "${NVKIND_SHA}"
(
cd /tmp/nvkind-src
GOFLAGS=-mod=mod go get "sigs.k8s.io/kind@${KIND_VERSION}"
GOFLAGS=-mod=mod go mod tidy
go mod vendor
go install ./cmd/nvkind
)

# `sg docker -c` avoids re-login after usermod -aG docker.
if sg docker -c "kind get clusters" | grep -qx "${NVKIND_CLUSTER_NAME}"; then
Expand Down
Loading