That sounds like an enormously expensive exercise.

As someone who's done something similar (https://blog.lukesalamone.com/posts/creating-tiny-semantic-s...) the expensive part wasn't the training itself but the data curation and evaluation post-training. For this, getting a reasonable distribution of tool calls when the tool call can be anything isn't easy.

Once you have that, the model is small enough batch sizes are probably enormous and training can probably be done on a consumer-grade GPU in a week or less. Or even faster on a bigger GPU.

At <50M parameters, training costs are completely trivial. You'll spend a lot more on your rent this month.

Haha, my rent is cheap lol

Training scales pretty badly, so smaller models like this are really not that bad in terms of cost.

At this size, it certainly doesn’t have to be.

You can train a model of this size on your laptop in a day.

So is the challenge, now, getting the right sort of data?