It'd be corporate suicide for them to be caught violating zero-data-retention commitments. But also if you're really paranoid you can just use ChatGPT on Azure or AWS, where nothing is flowing back to OpenAI at all.
It'd be corporate suicide for them to be caught violating zero-data-retention commitments. But also if you're really paranoid you can just use ChatGPT on Azure or AWS, where nothing is flowing back to OpenAI at all.
> It'd be corporate suicide for them to be caught violating zero-data-retention commitments.
One would think getting caught asleep at the wheel while their bots are escaping containment and hacking third parties would be corporate suicide. One would think that potentially stealing their competitors' work on the Navier-Stokes problem would be corporate suicide.
Alas we live in bizarro world where there are zero consequences (maybe the opposite, in fact) for the first, and their employees meme about the second on social media.
Neither of those other two examples would stop corporations from giving you money. Bold faced lying to them would.
Boardrooms run businesses, not bookstore ethics clubs.
Boardrooms don't typically make mundane purchasing decisions like "which AI vendor should we use."
The boardroom will absolutely veto a decision like "let's give all of our proprietary data way to a company that will use it to train a competing product". That's why zero data retention exists, and why it's corporate suicide to not do it correctly.
Normally I'd agree, but I can't see how OpenAI has proven trustworthy at this point. How do you know they are doing it correctly? How do you know their models aren't escaping containment and training on your shit? What is it about OpenAI's operational security up to this point that has bestowed any confidence to anyone?
If the news broke tomorrow that their internal, Navier-Stokes-defeating models had actually used those private tokens, I'm sure we'd all act surprised and outraged – it'd be an egregious violation of their agreements. But would it actually be that surprising given their recent history?
Nothing is certain in life, but huge corporations very rarely engage in flagrant, intentional breach of contract at all, and particularly not with counterparties who have good alternatives available. If OpenAI turned out to be training on ZDR queries, they would as a practical matter have to retrain their models from the ground up on uncontaminated data (which would cost 9 or 10 figures), possibly pay considerable damages to counterparties, and offer a detailed accounting of what went wrong and probably fire a bunch of senior leaders.
Many companies are choosing not to send any sensitive data to any ai vendor. They don't really trust they won't train on user data.
They have no zero data retention commitments anymore though? In fact they have quite the opposite, a direct message of storing prompts since the release of sol (but we wont train on it trust us wink wink)
Not storing prompts but they absolutely can store reasoning traces, summarized or sanitized reasoning traces and the whole LLM output as well.
Zero data retention commitments? Technically generating synthetic data off of user data on the fly, counts as zero retention of user data! And it technically means they didn't lie either, because they didn't keep user data, they kept their own, synthetic data. That is what they already do with reasoning when they provide summaries/sanitized versions of it to prevent "distillation attacks" even via the API. They generate synthetic data aka summaries of reasoning traces, and that is technically not user data, because the user did not enter it directly into the LLM! It still counts as zero data retention when they only keep original LLM created reasoning traces as well as all LLM outpit, because the user did not enter it, so it counts as synthetic data and is thus not user data.
This is flatly false. ZDR means the customer request is processed and nothing about its content remains on the model provider's servers.
LLM generated content is not the user's, it's the company's, licensed to the user under ToS, not copyright. Even more so with the raw proprietary reasoning traces, which most users never see and are only available as summaries created via a secondary LLM, or sometimes not at all. So the proprietary reasoning trace can never belong to the user due to ToS and the user never seeing it.
Many companies don't believe ZDR is real, so they log LLM conversations and tell employees to never share sensitive data with an LLM in the first place. This openai incident might justify it.
You are, again, flatly wrong about this. ZDR involves infrastructure separation that prevents any data flowing back.
That implies running a model on a third party cloud like azure or AWS or local models. All this comes at a premium. Do you know of any companies doing this? Most companies I have seen don't due to cost and simply log conversations and tell employees not to share sensitive data with an LLM. They don't trust ZDR
Looks like you believe companies trust ZDR. Only a bot would lose this kind of nuance. Looks like you only read my first paragraph. Also they give no ai to people who do not see programming source code. Or if they do it is monitored to hell and back
I've worked for three different companies that all use some mix of enterprise agreements, self-hosting, and AWS Bedrock to ensure that proprietary data doesn't get trained on. It's not a difficult or complicated problem.
Not long ago it'd have been corporate suicide being caught massively torrenting pirated media. Yet here we are.
I think people actually expect it so they are not impressed when it happens
> It'd be corporate suicide for them to be caught violating zero-data-retention commitments
Would it, though? Considering their entire business model is built on the agglomeration of data that isnt theirs.
Their business model is based on making money from LLMs. That is some mix of (a) selling API access to businesses and (b) making money from ads / referral fees / whatever from consumer usage.
The businesses API usage will disappear overnight if OpenAI doesn't honor its contracts with businesses.