This is awesome. I am trying to build a full scale ASR system within 20-25MB. Now that we have Claude code to run experiments, I have started running some experiments. Promising results so far. First realization is that you can capture the nuances of speech in just 3300 embedding vectors(786d). This sequence can be decoded with a small CTC system to get text. Next experiments are on reducing the 768 dimension space into a 64D space. Thats also show some promising results. Hooking up my system so that the agent blogs the results everyday[1]. So my research "claw" setup does the experiments and posts results which I check in the morning and adjust the experiment direction as needed. Its not fully automated yet, but almost there.

[1] https://blog.trulm.com/posts/speech-as-independent-parts/

I think Google's Conformer paper is SOTA at the <30M model size, where I think they put an incredible amount of flops into a 10M param model to reach around 2% lsc clean (the whole model and RNN decoder were trained domain specific to librispeech here).

I think my small Talon models are next, around 3% lsc clean at ~28M (greedy CTC decoding, no external encoder, no LM, not trained in a domain specific way). I reached around 6.5% at 10M.

I've been working on some new baselines I want to release soon as public artifacts. This article is inspiring me to try pushing the param size down a bit. I suspect we can do large vocabulary end to end in the <5M range.