Very cool, are you also testing local, cpu-runnable/trainable models/classifiers?

I trained some to do some interesting things, including playing doom: https://github.com/nicobrenner/jeffy

I’ll try training one for this benchmark, seems like fun

Not yet but the repo is open source and set up so that you can benchmark your own model and we can add it to the leaderboard, would be cool to see if Jeffy can play: https://github.com/opper-ai/jevman-benchmark/blob/main/CONTR...

Jev is useful to run locally overnight: it can classify the results while you sleep, ready for you to review in the morning.

Yes, try with toxic/toxichat and CFPB