The chatjimmy demo is using a model that needs 6-18GB of VRAM. That's not exactly trivial.
I could see it being feasible to get a Qwen-3.6-27b type of model done on something like this. Qwen-3.6-27b at 18tok/s would be a game changer.
The chatjimmy demo is using a model that needs 6-18GB of VRAM. That's not exactly trivial.
I could see it being feasible to get a Qwen-3.6-27b type of model done on something like this. Qwen-3.6-27b at 18tok/s would be a game changer.
right, but that's a reticle size chip. to put something in a phone it has to be ~10-30x smaller