Doesn't Ollama use llama.cpp so their point stands even if they used it directly?
It wouldn't be the first time ollama's llama.cpp fork reintroduced bugs and was missing important optimizations.
It wouldn't be the first time ollama's llama.cpp fork reintroduced bugs and was missing important optimizations.