I'm not really convinced that there's much secret sauce here, all the methods and data are public, the only real difference is how much compute it takes.
I'm not really convinced that there's much secret sauce here, all the methods and data are public, the only real difference is how much compute it takes.
everybody do be cooking with water. Chinese Labs provided pretty good, primarily cost-reducing techniques, like the sparse attention patterns recently. I'd bet OpenAI and Anthropic use their variants of those too, so they can get greater margin on their tokens - not something they'd really want to / need to self-report.