Abstract
Enterprises increasingly run GPU-bound workloads such as Large Language Model (LLM) inference inside virtualized infrastructure, yet practitioners have little systematic guidance on how the choice of virtualization platform affects GPU passthrough in configuration effort as much as performance. This study’s primary contribution is a structured, paired comparison of configuration complexity and performance across six platforms (Proxmox VM and LXC, native KVM, OpenStack VM and Zun, and Podman) on a single NVIDIA RTX 4500 Ada GPU, emphasizing the under-documented container paths (OpenStack Zun and Podman with the Container Device Interface). The platforms differ most in configuration complexity, where OpenStack, especially Zun, proved the most demanding and Proxmox and LXC the most straightforward. As a confirmatory baseline, for a small-model, low-concurrency vLLM workload that does not saturate the device, exclusive passthrough yields statistically equivalent throughput, duration, and inter-token latency across all platforms, within 2% of bare metal; this equivalence is bounded to the non-saturating regime tested and is not a general claim. Running two containers concurrently (each capped at 45% of GPU memory, compute time-sliced) doubled the offered load and raised aggregate throughput by 12–14%, consistent with idle GPU capacity but not isolating a GPU-sharing benefit.
IPC Classification
Keywords
€ 4.00