We had processor level branch predictors. Now do we not only pre fill the next prompt, why not just start generating the response as well?

Interesting thought at least.

[dead]