The claim that these models "run on almost any phone" as well as the in-app indicator of performance that says it may "run comfortably" are some serious exaggeration. Tried running Qwen 0.5B on mine [1]. It took five minutes to produce "I am a large language model created by Anthropic" (lmao) and then stopped outputting anything at all. I am not expecting you to somehow make the models perform better or whatnot, I just believe that the performance claims need review.
[1]: https://m.gsmarena.com/xiaomi_redmi_note_13_pro-12581.php
Alright, I have since experimented with the other harnesses mentioned in this thread (AI Edge and PocketPal) and both run the same models much, much faster. Gemma 4 E2B is extremely usable, Qwen simply flies. Same device. I don't know what you are doing, but it's clearly not great...
[dead]