I’m working on this problem using a vocab-free, byte-based approach. It’s definitely solvable.

https://huggingface.co/posts/omarkamali/593639295164067

https://huggingface.co/blog/omarkamali/tokenization

Careful: the problem is very certainly ___not___ counting letters. That is only a telling way to check "is the NN checking or not?". We demand that NNs for consultancy tasks check, strictly.

I used to think byte level tokenization was the answer, but humans also think at a word level and only reevaluate the words at a character level when asked. The solution to better tokenization across languages is likely to be learned tokenization. Here is one attempt I have seen: https://github.com/SamD770/bitter-lesson-tokenization