Hacker News

thewebguyd a day ago [ - ]

I'd go for at least 32GB+. It'll fit in 24GB but leaves you little to no room for context, and that's at 4-bit quantization.

If you want to run unquantized, you definitely need 128GB.

Catloafdev a day ago [ - ]

Nobody runs unquantized, there's literally no reason to. Q8 would be the largest anyone actually runs on consumer hardware for inference.

a day ago [ - ]

[deleted]

bityard a day ago [ - ]

Halving the precision of the weights is not a free lunch...

Catloafdev 20 hours ago [ - ]

Q8 is virtually lossless. The quantization is much more noticeable around Q4 and below. FP16->Q8 on consumer hardware is 2x the speed at ~99.99% the quality.

rvba 12 hours ago [ - ]

Any source that confirms the 99.99% quality?

bitexploder a day ago [ - ]

It also comes down to inference speed, not "can I run this". 8-bit quant is quite a bit slower on an M5 Pro.

gchamonlive a day ago [ - ]

[dead]