I think Google's Conformer paper is SOTA at the <30M model size, where I think they put an incredible amount of flops into a 10M param model to reach around 2% lsc clean (the whole model and RNN decoder were trained domain specific to librispeech here).

I think my small Talon models are next, around 3% lsc clean at ~28M (greedy CTC decoding, no external encoder, no LM, not trained in a domain specific way). I reached around 6.5% at 10M.

I've been working on some new baselines I want to release soon as public artifacts. This article is inspiring me to try pushing the param size down a bit. I suspect we can do large vocabulary end to end in the <5M range.