To some degree even unintentionally this is all inevitable. I know this isn’t what you’re taking about, but it’s a similar line of thought. These models are consuming public free information, and they’re all also producing free public information, so they’re all pissing and drinking into same pool.
If we go down the line of dead internet theory which I’m becoming more convinced of these days, the volume of information that’s not necessarily original or extracted from reality and just interpolated and extrapolated from existing information in different ways by LLMs will greatly outnumber human information coming up.
In which case these models should.. start to converge on the same data I imagine, with slightly different behaviors within that. One big generative orgy feedback loop.
Many of the Chinese optimizations are public because they are published by the Chinese labs themselves. Hard to say what the labs are doing of course.
DeepSeek has published some really good papers. Lately they're pushing really hard for dramatically cheaper serving costs.
I may be remembering wrong, but reasoning was first demonstrated by DeepSeek. Edit: I am indeed remembering wrong, seems o1 was first.
Deepseek’s R1 paper was the first paper to describe how to do it. O1 was released before R1 came out.
A month before the R1 paper came out, they released the Deepseek math paper which described their method for MoE load balancing.
To some degree even unintentionally this is all inevitable. I know this isn’t what you’re taking about, but it’s a similar line of thought. These models are consuming public free information, and they’re all also producing free public information, so they’re all pissing and drinking into same pool.
If we go down the line of dead internet theory which I’m becoming more convinced of these days, the volume of information that’s not necessarily original or extracted from reality and just interpolated and extrapolated from existing information in different ways by LLMs will greatly outnumber human information coming up.
In which case these models should.. start to converge on the same data I imagine, with slightly different behaviors within that. One big generative orgy feedback loop.
Let me put it this way: if they're not ripping off all that they can, their investors need to shitcan the leadership.