So YES if you're using Cowork or Chat

No, because they are cached, the inference cost is paid once per model, does not scale linearly per user or use.