Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -258,6 +258,12 @@ To create the config map:
kubectl create cm <config map name> --from-file=config=alumet-agent-client.toml
```

### Using the nvml plugin

When enabling the nvml plugin through the values, the chart automatically appends a limits of `nvidia.com/gpu` to `1`.
You can override the limit value in the chart's values.
Comment thread
titouanj marked this conversation as resolved.
Consider using [Time Slicing](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/gpu-sharing.html) to enable GPU sharing.

### Configure runtimeClassName

If you're running ALUMET in a specific environment: with nvidia GPUs, with multiple runtimes, etc.
Expand Down
4 changes: 2 additions & 2 deletions charts/alumet/Chart.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ type: application
# This is the chart version. This version number should be incremented each time you make changes
# to the chart and its templates, including the app version.
# Versions are expected to follow Semantic Versioning (https://semver.org/)
version: 0.6.2
version: 0.6.3

# This is the version number of the application being deployed. This version number should be
# incremented each time you make changes to the application. Versions are not expected to
Expand All @@ -33,5 +33,5 @@ dependencies:
version: 0.6.2
condition: alumet-relay-server.enabled
- name: alumet-relay-client
version: 0.6.2
version: 0.6.3
condition: alumet-relay-client.enabled
2 changes: 1 addition & 1 deletion charts/alumet/charts/alumet-relay-client/Chart.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ type: application
# This is the chart version. This version number should be incremented each time you make changes
# to the chart and its templates, including the app version.
# Versions are expected to follow Semantic Versioning (https://semver.org/)
version: 0.6.2
version: 0.6.3

# This is the version number of the application being deployed. This version number should be
# incremented each time you make changes to the application. Versions are not expected to
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,14 @@ spec:
imagePullPolicy: IfNotPresent
resources:
{{- toYaml .Values.resources | nindent 12 }}
{{- if .Values.plugins.nvml.enable -}}
{{- if empty .Values.resources.limits }}
limits:
{{- end }}
{{- if eq (index .Values.resources.limits "nvidia.com/gpu") nil }}
nvidia.com/gpu: 1
{{- end }}
{{- end }}
volumeMounts:
- name: sysfs
mountPath: /sys
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -96,4 +96,45 @@ tests:
path: spec.template.spec.containers[0].env
content:
name: RUST_BACKTRACE
value: ""
value: ""
- it: shouldn't have limit gpu if nvml plugin is not enable
set:
plugins:
nvml:
enable: false
asserts:
- notExists:
path: spec.template.spec.containers[0].resources.limits["nvidia.com/gpu"]
- it: should have limit gpu if nvml plugin is enable
set:
plugins:
nvml:
enable: true
asserts:
- isSubset:
path: spec.template.spec.containers[0].resources.limits
content:
nvidia.com/gpu: 1
- equal:
path: spec.template.spec.containers[0].resources.limits["nvidia.com/gpu"]
value: 1
- it: should have limit gpu only once when nvml enable and limit set manually
set:
plugins:
nvml:
enable: true
resources:
limits:
nvidia.com/gpu: 32
asserts:
- isSubset:
path: spec.template.spec.containers[0].resources.limits
content:
nvidia.com/gpu: 32
- isNotSubset:
path: spec.template.spec.containers[0].resources.limits
content:
nvidia.com/gpu: 1
- equal:
path: spec.template.spec.containers[0].resources.limits["nvidia.com/gpu"]
value: 32
Loading