> https://github.com/darshi1337/apogee/blob/main/MODELS.md#loc...

> Unlike Ollama, [llama.cpp] serves one model at a time: the GGUF you launched it with.

FWIW: llama-server's supported multiple models for a while now:

https://github.com/ggml-org/llama.cpp/tree/master/tools/serv...

Ohh thank you for your suggestion. I will try implementing it this weekend.