This is awesome. I will take some time to dig in. When I am not working for client(s), I focus entirely on tiny LLMs - I have specific approach to prompting, avoid multi-turn chat and build harness to fit the selected LLM as closely as possible.

My experiments are in https://github.com/brainless/

I will be happy to share what I learn.