Anyone else somewhat weirded by current state of affairs that this sort of information is valuable enough to even bother selling... And that it actually happens... It feels like some societies are in really weird place.

How does this have value? Is any and every sentence in an e-mail considered 'fact' and thus to be fed into the AI?

90% of e-mails and Teams communications are inane. Polite banter, "thanks for taking care of that, I appreciate it" "please route the forms to Janet this week because Bill is on vacation" "unit will be un available until the parts come in" . I can't see the intrinsic fact value of this kind of communication without screening it. And after screening, the gold nuggets would be minimal.

> How does this have value? Is any and every sentence in an e-mail considered 'fact' and thus to be fed into the AI?

I have a hunch what this is for. AI companies want to make bigger inroads into nontechnical work settings. But LLM progress outside of fields where verifiable rewards for RL post-training can be synthetically generated (coding, math) has been pretty flat. Buying years of operational data from a company like an airline could be used to reconstruct long-horizon task trajectories in areas like customer service or marketing.

> I have a hunch what this is for. AI companies want to make bigger inroads into nontechnical work settings. But LLM progress outside of fields where verifiable rewards for RL post-training can be synthetically generated (coding, math) has been pretty flat. Buying years of operational data from a company like an airline could be used to reconstruct long-horizon task trajectories in areas like customer service or marketing.

This. Training data companies like Mercor are even creating simulated companies to create similar data, so they can better automate even more white collar work:

https://www.nytimes.com/2026/07/10/business/ai-white-collar-...:

> The data-training start-ups see a lucrative opportunity in recreating workplaces in miniature: controlled environments in which their gig workers can evaluate and reproduce emails, memos and slide presentations in context. The information emerging from such a setup, the companies boast, will help shrink the gap between what A.I. models can accomplish and what office workers actually do from one minute to the next, as ideas and instructions flow between meetings, documents and applications.

> Scale, for example, has said that “our environments replicate real-world workflows,” and that its contractors “curate artifacts that capture the complexity, ambiguity and edge cases of real professional work.”

> Executives see the models’ shortcomings as a sign there’s more for them to do.

> “I often use Claude Cowork, right?” said Edwin Chen, the founder of Surge. “And even though Claude Cowork is incredibly smart, oftentimes it doesn’t quite understand the nuance of Slack. It doesn’t quite understand this ambiguous question I have. It doesn’t quite understand where to go and find this Google document.”

> According to the Bloomberg Billionaires Index, Mr. Chen’s stake in Surge — and his vision for what it could become — makes him the 258th-richest person in the world.

> “I often think about us as essentially, like, the school for A.G.I.,” Mr. Chen said, referring to a prophesied level of A.I. that surpasses human intelligence. “A.I. comes to us, and A.I. learns to run the world.” He and his customers at the big A.I. labs, he said, are designing the curriculum.

LLMs aren't a database. They're an attempt at brute-forcing an artificial mind. The who and what aren't really interesting there, it'll forget most of such details anyway. What matters is the patterns visible in the text at various scales. How people write. Why they write. To whom they write, in response to what. How does e-mails about mistakes correlate with PDFs they're referring to. How people work with ticketing systems - like how, actually, a ticket plays out. The jargon, the acronyms, the vibes, the causal links. It's all in there, and it's another slice through the set of things humans do, to be combined with other slices already in the training data, and enriching the whole.

(Something something we will add your distinctiveness to our own, you will be assimilated, ...)

(Hell, the fact that it's all from one org would make it a great dataset to have in the open for sociological studies. I bet that today, aided by LLMs to sift through it, you could use it to map how information flows through a large org - how incident on the floor travels through time and layers of management until it reaches the C-suite, what of it survives, how it gets reacted to, how the reactions flow down...)

A major selling point of AI chat bots is for customer service automation. A clean dataset like this is a huge find. I'm not sure what you mean by "facts" or "gold nuggets". Training data doesn't need to be factual.

It will have a different writing style from the average blog post. Maybe they just want to train an AI that sounds less like an AI. Or one that speaks in vapid management style.

I am very much weirder out by it, yeah. Seems some societies are just excessively desperate for some kind, any kind, of fuel for economic growth, to the point this is where attention is now. The term "post capitalism" being thrown around feels less ridiculous than it did in years gone past.

Truth is even weirder. plenty of growth is possible but modern liberal democracy requires that all progress must be contingent on the production of enormous amounts of text that nobody would ever read.

3 years ago, the Wall Street journal covered a company trying to use AI to generate documents required for the approval of new nuclear reactor reactor designs, which sounds dangerous, until you get to the point where they'd cite 2 million pages as necessary for a typical application. [1] I don't need to explain why no individual or institution could read that, much less examine it in detail. I remember a rather funny question I found in a comment to that story - "How many pages of those 2 million could contain pornographic images before anyone notices?"

It's obvious why companies are so desperate for training data - a text generator of sufficient quality is more conductive to the growth of the nuclear industry than any scientific breakthrough in nuclear physics. (And of course, if you want to prevent the development of a nuclear reactor by your competitors, being able to produce millions of pages of high-quality objections will do the trick.)

And it's not just nuclear power. When it comes to stuff like building rail lines, apartments or power plants (both conventional and renewable), you'd find that the main bottleneck is the necessity to produce documents. And of course, many documents can be subject to judicial review - a process that consumes even more text.

[1] https://www.wsj.com/tech/ai/microsoft-targets-nuclear-to-pow...

Any kind of fuel for giving active investors that FOMO tingle which then forces the steamroll of index funds to blindly follow.

I guess the appropriation "any sufficiently advanced stock market is indistinguishable from entertainment" doesn't quite stop at equating the trade floor with a casino. At some point, entertainment also becomes the modus operandi of corporations.

Companies buying other companies files have been a thing since companies.