Could've been better if GGUF implemented QTIP format. GGUF representation is a major limitation for llama.cpp quantization performance

They use their own llama fork anyway, so that shouldn't matter.