
News & Updates
8 min
Two RTX 4090s Over PCIe: Splitting a 27B Model Beat Two Replicas by 2 to 4x
We had a rule that a model which fits on one GPU should run as one replica per GPU. On a two-RTX-4090 machine with a 27B model, that rule was wrong by a factor of two to four, and the reason was not the parallelism at all. Here are the measurements, the rule that replaced it, and how the router now sends a 40k-token prompt to the worker that can hold it.
inferencevllmbenchmarks