To some degree even unintentionally this is all inevitable. I know this isn’t what you’re taking about, but it’s a similar line of thought. These models are consuming public free information, and they’re all also producing free public information, so they’re all pissing and drinking into same pool.
If we go down the line of dead internet theory which I’m becoming more convinced of these days, the volume of information that’s not necessarily original or extracted from reality and just interpolated and extrapolated from existing information in different ways by LLMs will greatly outnumber human information coming up.
In which case these models should.. start to converge on the same data I imagine, with slightly different behaviors within that. One big generative orgy feedback loop.