The other day I saw a benchmark of coding models that fit in 8GB VRAM - for reference some version of Mistral was added, normally requiring 32GB, but moving along at 4-5tok/sec when partially offloaded to CPU.

Surprisingly, some of the small models would not only give worse results, but also took longer than Mistral, because they were thinking so much.

That is an important detail which I was previously overlooking.