I'm not sure what this means for AI startups if their innovations can be copied by OSS so quickly (what, like 2 weeks?). There's "consumer surplus" for everyone, to borrow an economic concept. But we do ideally want some of the surplus to flow to the innovator, too. I know there were precursors, but that's fine - it's hard to have a totally novel idea in such a popular field. I don't know what the end game is for TypeSafe - they'd need to demonstrate perpetually better results, or compete in another axis: UX, support, custom solutions, etc. So much of the time, someone proving a concept, or it simply getting enough publicity, is enough for a "Cambrian explosion" of follow-ups and copies. Famously, that was true for "Attention is All You Need", and the general idea of "next-token prediction" being so powerful.
We've stumbled into general differentiable models..
I think the idea of a "feature startup" is dead. What used to be a niche subscription business is now an individual Epic level of work. The smallest viable business becomes what two or three years ago was a mid tier enterprise. It is no longer "look at this tool I maintain", but "we take this specific approach using these hundreds of tools merged together to solve a problem in a specific way that nobody is going to compete with. Not because they can't compete if they wanted to, but that the competitions approach diverges in fifty different chosen ways that they are targeting a different market segment essentially."
I adhere to the idea that this is software's "Tower of Babel" moment where everyone just fundamentally ships things in completely diverging architectures, because creating a ground up architecture is no longer something that needs to be avoided for an economically viable business mode that in the past two decades would have otherwise incentivized people into industry standards. In a world where "taste" is the focus, single ingredients in the recipe aren't enough.
Because what they did is kinda trivial. Its basically like the Dropbox comment really[0], except here you don't need petabytes of storage and infinite VC pockets.
After chatgpt everything in AI mostly became LLMs and building wrappers around them. It's like people forgot how to do ML.
To those of us who actually trained models back in the day, its kind of cute to see people wowed by a classifier. Yes, this is 0 shot and doesn't need training (most people wanting this would've used structured output, this is cool because it's cheaper and faster). But anyone with basic ML knowledge could've built this in a few hours.
The question is mostly why wasn't this productized. And it's interesting indeed that it took this long to become a finished product.
[0] https://news.ycombinator.com/item?id=9224
> The question is mostly why wasn't this productized.
Probably because doing it wrong (using an llm in place of a classifier) is more profitable? (For the people selling inference.)
Today we might need to evaluate “innovation” in a new standard, and have a different expectation for what innovator would be awarded. Getting public attention in such an era where innovation happens every a few days could’ve already been something precious. And that attention would allow TypeSafe to be heard easily next time. Like OpenAI, Anthropic, or any others, they launch frequently but still each time they launch something new, that would hit headlines. I think that’s the “surplus” flown to innovators today.
This is precisely why we need strict government regulation of open weight models. :wink:
Are you saying laya copied from jev, and released in two weeks? If so I don’t thinks it’s quite as simple a story as that. https://xtxinversexty.com/layas-prior-art-claim-is-absurd/
It a paradox when the article is claiming the prior art is absurd, but then goes on to analyse the one side and compare it to another for which most of the values (except scaling the concept) are unknown. And even for scaling, it uses the first, pre-laya instance to judge the limited schema, while overlooking that Laya is just doing this scaling. Important to note that prior art is not having built the exact same thing.
The moat is the RL synthetic data pipeline they set up to train jev. Open sourcing that would be the coup, not the model architecture and training scripts, which are trivial.
Presumably the training recipe and training dataset itself cannot be easily copied in a week or two. So if they want to shut down these competitor models they need to make it obvious how they are better than them.
[dead]