I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.
I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.
How do you cleanse the data at this scale?
By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data.
We have a lot of details in the tech report if you want to go deeper.
Hey! I’m curious if you tried comparing luxical to model2vec classifiers for the pretraining.
I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!
Are there plans to make much larger versions of this model? With 500B-1T params for general purpose knowledge tasks, similar to the current leading proprietary models?
Stay tuned, this is just the beginning. Merging with Cohere will help in scaling, too.
How exactly would an open project do that?
How much should German authors be paid, $3,000 per book, like Anthropic paid?
Why expose yourself to this liability?