I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting:
> Improve the model for everyone
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.
It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.
I refer you to this:
https://news.ycombinator.com/item?id=49643556
Quoting:
> "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."
Old Facebook trick - likely resetting that box each time the app is updated.
"We've made some updates to improve the security and privacy experience" => "We've changed some of the options available and reset everyone to defaults"
In all fairness this is someone saying something. Misremembering happens. Unless we have something with a bit more evidence, the simpler explanation suffices.
In more accurate assessment rather than assuming no maliciousness nor incompetence, remembering also happens. Unless we fail as a society, the simpler explanation that "OpenAI is training on all data it can and resetting config toggles because it uses the same cohort of engineers that came from Meta and other FANGAMAAMMAM clones" suffices.
I wasn't aware of that, definitely a shady practice if that's the case.
>"reset this more than once ... to my surprise I found it re-enabled"
At least when my computer's bluetooth exhibits this behavior (e.g: if you don't have a keyboard&mouse plugged in at boot, bluetooth might auto-enable), I can go inside the hardware and physically disconnect the antenna.
What am I supposed to do in software (perhaps hardcode config.file)? in cloud software services (??)?
Has anyone else seen this happen? I checked and my setting is still off.
There's a difference between "this is allowed under their ToS" and "it is academically unethical to fail to credit the people whose specific conversations were fed into a model that was used to solve a problem".
I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).
Whoa that's a slippery slope! Next you'll want model runners to cite the data their models were trained on
In fact we should though.
I suspect the fundamental problem here is it's hard (if not impossible) to determine if someone who tried the winning approach deserves the credit for the discovery, because there's always the chance that they could've done something differently, or stopped before finishing, and thus never actually made the discovery. They might've even tried the approach just based on a whim, without really thinking it would work, and might've given up without a final insight. And fundings run out, people end up in hospitals, etc. What do you credit them with when the work isn't finished? For trying an approach that sounded promising? You can do that I guess, but is that what they want?
But who gets credit then? Every mathematician who's work was read by an LLM during training? By that logic, we should put every published mathematician's name on the authorship of this paper. Sure, this guy should be higher up the list, but everyone's name should be on it by standard academic convention.
But this gets back to the original "who owns the LLM output" and "can you train models on the internet" argument that's been raging for years.
Does every mathematician get cited in every maths paper? I think it's pretty clear who should be cited.
Not unticking a box in settings doesn't constitute consent in my opinion. I'd never put anything I value into ChatGPT anyway, though.
Under EU rules it doesn't constitute consent.
"Improve the model for everyone" can be implemented in so many ambiguous ways.
https://news.ycombinator.com/item?id=49643513
>>"Improve the model for everyone"
e.g: allows us to sell your personal data to make money so we can continue offering this service to all customers
I'm done with weasle-words and hours-long EULAs – we're at the point where USA needs to catch up to EU's consumer protections, perhaps with laws similar to already-existing USA "truth in lending" requirements (e.g: interest rates must be prominently displayed in a larger font, including annual fees, on all credit offers).
----
My judge-brother always asked during our childhood "why don't you think the judicial system is fair?!?" Thirty years ago, the best I could offer was "because it's a two-tiered system that mostly (only) rich people can afford to participate within."
Now my answer is: "the best example I can give is that our judicial system allows binding arbitration [and qualified immunity for police]. The system is set up so corporate personhood is more important than humanity, and it shows."
I believe it can be "off by default" depending on terms negotiated between the enterprise customer and ChatGPT.
We have ChatGPT at work and it explicitly says that "workspace data isn't used to train models"
> We take steps to protect your privacy
No mention what those “steps” are, success criteria, or whether they are successful by any navies at all … they take steps though… so it’s fine, and if we know one thing it’s that we can really truly trust someone off the likes of Sam Altman.
Ok, you shut it down, or that is what they make you believe. You give the instruction to shut down, you can't know if it has been applied.
Yes but what about “analytical purposes” what does that cover and can you turn it off? I have found out you cannot. It’s the Trojan backdoor to your data.