GPU sharing via Timeslicing in Kubernetes

#writeup #GPUs #low-level

During my time at IBM, I've had the opportunity to work on improving the performance of our self-hosted models running on our GPU clusters, and as usual when I dive into a new project, I go down a lot of rabbit holes... One of these being GPU time-slicing.

Your home PC likely has an NVIDIA GPU that gets requested by your operating system and drivers to run graphics-intensive applications, but what if you want to share your GPU with another computer?

This probably sounds weird, but it's really important to be able to do this, especially in the world of AI inference.