Not today, but Interesting idea. The main motivation was a drop-in for an existing Jev setup, so the only teacher right now is Jev and the audit measures agreement with Jev. A correction would have to become a second label source that overrides Jev's for that input.
The head is a multinomial logistic regression: one linear layer plus softmax on top of a frozen sentence-embedding model (bge-small by default, swappable). That head is the entire local model, the encoder is off the shelf and never changes.
Yeah, I kinda figured that head was the whole model, the way you phrased it. scikit-learn?
No, plain numpy. It's full-batch Adam on cross-entropy against soft targets, about 80 lines.
I'll look into Adam, thank you!