The value of an LLM is the dynamic reasoning you get out of it and the cost to execute on that.
I see two forces working against this that proprietary models will always have over an open source model.
1. The biggest is content licensing. Content is quickly becoming gated by systems at the front of their load balancers, completely changing the social contract of the Internet. What used to be a quick google search for recent facts that lead me to places like reddit or twitter, is now completely walled off if you're not physically at your browser and using an IP address from a last-mile provider.
LLMs have pre-trained on the bulk of the information up to 2024/2025, but over time that will be more and more out of date.
Anthropic, OpenAI and Google will all have to pay for access to a lot of this content refresh going forward, and it does make a material difference in the output you get.
2. Liability is the other. A corporation can look at a contract for model access and see one that provides uptime guarentees, content infringement promises and model safety, and pick the contract that shields the corporation from the most liability. A 3rd party hosting platform like fireworks.ai that hosts open weights models won't provide any of that at all. They will simply bill you for time spent on their hardware and make promises that they won't log or inspect corporate traffic.
> A 3rd party hosting platform like fireworks.ai that hosts open weights models won't provide any of that at all.
Why couldn't they?