It can’t possibly be more expensive than doing equivalent work using context instead of RAM and inference instead of CPU.
Unless we’ve got the wrong balance of compute availability and inference availability right now, but I would expect the market to stabilize at some point.
It can’t possibly be more expensive than doing equivalent work using context instead of RAM and inference instead of CPU.
Unless we’ve got the wrong balance of compute availability and inference availability right now, but I would expect the market to stabilize at some point.