Those cards should be able to do a lot more. They support int8 and they have very good processing speed. For reference I have a single V100 running around 900t/s prefill and 100 t/s decode during DeepSWE runs on 27B (4 bit quant). Those cards have half the bandwidth so a realistic decode is going to stop around 50 t/s, but those cards can support NVFP4 more easily than the V100. Combined with a TP build you should be able to get 100 t/s. And you should be able to 2x the prefill of a V100. I have had Claude optimizing llama.cpp for about a week and am at around 300 t/s decode so far on two ancient V100s @ 64 GB RAM.
Also for the 3 people that ever read this and are curious about local models still, Qwen 27B 3.8 matched Sonnet 5 in the 17 DeepSWE tasks I have run so far, solving the exact same 7 it has. Caveat: datacurve combined low/medium/high Sonnet 5 data.
Those cards should be able to do a lot more. They support int8 and they have very good processing speed. For reference I have a single V100 running around 900t/s prefill and 100 t/s decode during DeepSWE runs on 27B (4 bit quant). Those cards have half the bandwidth so a realistic decode is going to stop around 50 t/s, but those cards can support NVFP4 more easily than the V100. Combined with a TP build you should be able to get 100 t/s. And you should be able to 2x the prefill of a V100. I have had Claude optimizing llama.cpp for about a week and am at around 300 t/s decode so far on two ancient V100s @ 64 GB RAM.
Also for the 3 people that ever read this and are curious about local models still, Qwen 27B 3.8 matched Sonnet 5 in the 17 DeepSWE tasks I have run so far, solving the exact same 7 it has. Caveat: datacurve combined low/medium/high Sonnet 5 data.
Very interesting! Will you be sharing the optimizations ?
I will make sure they get out there soon :)