At that point, how does this compare with simply running the model on the CPU?

Not an answer to your question, but maybe this has more info? I think with these optimizations and quants, it compares negatively. But can these optimizations be applied to models you want? Another question. https://news.ycombinator.com/item?id=48353348