llama.cpp sycl and vllm xmx work is pretty incredible right now - you just gotta build it with some extra flags
Llama would be nice for the ggufs. Any specific flags or tutorials I should look at?
The docs are a great start.
https://github.com/ggml-org/llama.cpp/blob/master/docs/backe...
Llama would be nice for the ggufs. Any specific flags or tutorials I should look at?
The docs are a great start.
https://github.com/ggml-org/llama.cpp/blob/master/docs/backe...