> I don't know if his solution (""We should all go insane building interlocking evaluation and optimization pipelines, instead.") would be the long-term solution.

TFA could explain this one part better I think. The whole process proposed is real with lots of stuff in the literature, but by definition NOT a long-term solution in the sense that this process actually has no end. None of the approaches can get you a static answer for a moving target/platform.

So the "interlocking pipelines" for eval/opt would not be some stepping stone you can throw away, and they aren't something you'd run periodically. They'd basically be always on forever and spending 10-100x on system complexity and on tokens. Unless of course you're ready to freeze everything else about the whole system forever (including the backend model, and the whole nature of the "average" context window, the plugins/other prompts in the mix, etc).

Are most people in position to freeze requirements/platform forever? Not really, because if they were they'd just build a fairly static system and probably have limited use for AI. Are most people in a position to just casually accept 100x complexity/cost? Not really, that's the "it's not yet webscale" kind of advice that sounds good but isn't necessarily reasonable for average use-case or average org. Since specializing your own locale for this is usually a mistake.. the likely future direction is eval/optimization as a service