Is that similar to what ggerganov is talking about here?

https://x.com/ggerganov/status/2089214161884414147

Yes - it would seem so :)

I did make a PR to the official llama-cpp repo some time back (about a month or so), but abandoned it as there seemed to be too much community concern that the mechanism would degrade model performance... Perhaps i'll polish it up and put some effort into benchmarking and revive the project in the near future.

would be great as an opt-in though