what character prediction rates are you getting on some unseen datasets?

The held-out scores reported in the Readme IS the unseen dataset.

These are not successful prediction rate per char though.

I got you, will add later to the repository.

Thanks, it seems like a nice way to compare effectiveness of different non-typical methods which are not yet capable of some more ambitious benchmarks.