Enable Mthreads GPU sharing
Introduction
HAMi now supports mthreads.com/vgpu by implementing most device-sharing features as NVIDIA GPUs, including:
GPU sharing: Each task can allocate a portion of GPU instead of a whole GPU card, thus GPU can be shared among multiple tasks.
Device Memory Control: GPUs can be allocated with a specific device memory size on certain types (e.g., MTT S4000, MTT S5000), with hard limits enforced to prevent exceeding the allocation.
Device Core Control: GPUs can be allocated with limited compute cores on certain types (e.g., MTT S4000, MTT S5000), with hard limits enforced to prevent exceeding the allocation.
Important Notes
-
Device sharing for multi-cards is not supported.
-
Only one Mthreads device can be shared in a pod (even if there are multiple containers).
-
Support allocating exclusive Mthreads GPU by specifying mthreads.com/vgpu only.
-
These features are tested on MTT S4000 and MTT S5000. On MTT S5000 clusters, set
devices.mthreads.memoryPerCardto[160]when installing HAMi, as shown in Enabling GPU-sharing Support.
Card specifications
Both card models expose 16 core groups per card. Device memory is requested in 512 MiB units, and the valid values depend on the card capacity:
| Card model | Device memory | Total sgpu-memory units | Valid sgpu-memory values |
|---|---|---|---|
| MTT S4000 | 48 GiB | 96 | 2, 4, 8, 16, 32, 64, 96 |
| MTT S5000 | 80 GiB | 160 | 2, 4, 8, 16, 32, 64, 128, 160 |
Requests with values outside the valid list are rejected by the admission webhook. The per-card capacity is controlled by the cluster-level devices.mthreads.memoryPerCard chart value, a list with one entry per card model. A mixed S4000/S5000 fleet is supported with [96, 160], so separate node pools are not required. Values above 96 per card additionally require a HAMi release with the memoryPerCard feature, see Use HAMi with Mthreads MTT S5000.
Prerequisites
- MT CloudNative Toolkits > 1.9.0
- driver version >= 1.2.0
- For MTT S5000 with sGPU: MT Container Toolkit >= 2.1.0 and MTML >= 2.1.0, with sGPU enabled through the Mthreads GPU Operator. See Use HAMi with Mthreads MTT S5000 for the full setup.
Enabling GPU-sharing Support
- Deploy MT-CloudNative Toolkit on Mthreads nodes (Please consult your device provider to acquire its package and document)
You can remove mt-mutating-webhook and mt-gpu-scheduler after installation (optional). HAMi's scheduler and webhook take over their roles. On MTT S5000 clusters running the Mthreads GPU Operator, disable the vendor components through the ClusterPolicy as described in the MTT S5000 installation guide.
- Set
devices.mthreads.enabled=truewhen installing HAMi. For a default install (MTT S4000, or any 96-unit card), run:
helm install hami hami-charts/hami --set scheduler.kubeScheduler.image.tag={your kubernetes version} --set devices.mthreads.enabled=true -n kube-system
- On MTT S5000 clusters, use a values file that also carries the kubeScheduler tag and the per-card memory capacity instead of the one-liner above:
# Set the scheduler image tag for your Kubernetes version.
scheduler:
kubeScheduler:
image:
tag: { your kubernetes version }
devices:
mthreads:
enabled: true
# MTT S5000 has 80 GiB device memory = 160 x 512 MiB units per card.
# The chart default (96) matches the MTT S4000 and must be overridden for S5000.
memoryPerCard:
- 160
helm install hami hami-charts/hami -n kube-system -f values.yaml
- If HAMi is already installed, apply the same values with
helm upgradeinstead of reinstalling:
helm upgrade hami hami-charts/hami -n kube-system -f values.yaml
The device config change is not rolled automatically; restart the scheduler after upgrading:
kubectl -n kube-system rollout restart deploy/hami-scheduler
Running Mthreads jobs
Mthreads GPUs can now be requested by a container using the mthreads.com/vgpu, mthreads.com/sgpu-memory and mthreads.com/sgpu-core resource type:
apiVersion: v1
kind: Pod
metadata:
name: gpushare-pod-default
spec:
restartPolicy: OnFailure
containers:
- image: core.harbor.zlidc.mthreads.com:30003/mt-ai/lm-qy2:v17-mpc
imagePullPolicy: IfNotPresent
name: gpushare-pod-1
command: ["sleep"]
args: ["100000"]
resources:
limits:
mthreads.com/vgpu: 1
mthreads.com/sgpu-memory: 32
mthreads.com/sgpu-core: 8
Each unit of mthreads.com/sgpu-memory represents 512 MiB of device memory. The example requests 32 units, or 16 GiB. Valid values per card model are listed in Card specifications. More examples are available in the examples/mthreads folder.