I am sympathetic but this buries the lede, hard.

You are competitive with Jev only if you fine tune on the train dataset and calibrate per question.

As much as I dislike literally everything about typesafes behavior, they have an API model that works on any problem without fine tuning, and that is the key.

To be honest everyone who can finetune can likely finetune a BERT for a specific task and get similar results to yours. And that has been true for years

The key to Jevs success is that it works without fine tuning