
News & Updates
8 min
The Faster Card Was Half the Speed
Two RTX 4090s on our network, same container, same model, same arguments. One runs Windows with WSL2 and benchmarks better on every hardware axis we measure. It serves inference 48% slower — three times the penalty the published figures report. Here is the measurement, what we ruled out, and why we think large quantized models are the worst case for GPU virtualisation.
engineeringbenchmarksdepin