Issue is that llama.cpp is the best way to run models on hardware that isn't nvidias.

Except when they have less than 16 gb of ram?