Who is going to upgrade the models ?

Who is going to fix it when the api does something weird ?

Who is going to proactively make sure it’s not overheating?

Chat GPT has enterprise contracts for a reason.

The same person managing the company's email accounts and whatnot.

I didn't say it's fire-and-forget. I'm saying all that is maybe a day of work every 3 months.

It doesn’t really work like that.

The companies which have the will and the budget to host their own LLMs typically require a ton of other, much smaller, models as well. They have internal security requirements, guardrails, audit, critical workflows start depending on your onprem setup, downtime is now something that’s not even allowed. There’s going to be a zoo of tooling, lots of bespoke work with internal clients who have no clue about docker, but now want their vibe-coded app to access a model they downloaded yesterday, and this model better be served and monitored 24/7 because now C-levels use it.

No one is going to budget several millions to then look at an email admin who hears ‘cuda’ for the first time in their life and ask them to just support the entire thing somehow.

And god forbid it’s an AMD setup.

This obviously isn’t relevant for a 10-person startup and their second-hand xeon with a single H100.

From my experience only the companies which are REALLY interested in privacy and data security bother with hosting their own models

My company is not even the same ballpark, but even we already have people for it. And doesn't include the fact that you can get a colo location and just a MSP or a contractor to do it for you

All of the answers to your questions above are in gp's comment already:

> A company of the size that this is worthwhile for, probably has dedicated devops on staff already

[deleted]