Couldn't they just grab and run an open weight model to save on API tokens?

You get better performance if you also finetune it for your task