Topic
Everything filed under AI Infrastructure, newest first.
RSS · JSON · All topics
GPU sharing via Timeslicing in Kubernetes
A write-up on NVIDIA GPU time-slicing in Kubernetes: how the GPU Operator advertises virtual replicas of one GPU so several pods can share it, the context-switching overhead and per-pod observability gaps this brings, how it compares with MPS and MIG, and a glossary of the key terms.
5 min · 1,232 words