I’d rather have one European AI lab like Mistral, with the financial firepower and compute to pretrain its own large models, than 20 weak labs that fine-tune Chinese models.
AI needs big bucks, and Mistral is Europe’s Anthropic.
I’d rather have one European AI lab like Mistral, with the financial firepower and compute to pretrain its own large models, than 20 weak labs that fine-tune Chinese models.
AI needs big bucks, and Mistral is Europe’s Anthropic.
just curious, how do we know that le chonk isn't just a fine-tuned chinese model? and/or distilled from US models?
It's in the announcement: "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe".
https://mistral.ai/news/mistral-large-4/
Once it's open weight people will be able to inspect and compare it's tokenizer, architecture etc and tell.
it's extremely unlikely that they re-use anything from a chinese model, that would be obvious quickly, what's more likely is using documents produced by a better model to create synthetic data.
It would not take them so long to train it. Their pace would be closer to the Chinese models.