More great work on local model but you’re still losing a lot. Down to 2 bit quantization and the coder model throws away half the MoE experts. In a world where anything is better than nothing, this is a net win. But we have a way to go still.

The real problem is the DRAM mafia and artificial scarcity. One of my notebooks is almost 3 years old, effectively similar spec now - same price (a bit higher actually). My desktop PC built around march/april 2023 (4090, 64gb ram, 7900x3d) is now pretty much still the top dog out there due to the gpu and fast ram insanity and if I wanted to sell it today, I'd get more money for it now used and over 3 years old that when I bought it!

We should be having 64/72+ GB video cards by now. 128GB+ system ram prosumer laptops and 256GB+ system ram prosumer/gamer desktops. But it all went to shit and it will require some brutal datacenter and datacenter-adjacent bankruptcies before it gets better.

Some of these greedy bastards need to lose their pants on all of this.

This is true, progress has stopped on the RAM and VRAM front.

Many new laptops come with 8GB as standard, the same as 12 years ago.

My 1060 from 2016 has 6GB of VRAM. A 5060 from 2025 has 8GB.

You can run IQ3_XXS, IQ3_S and the IQ4_XS quants on this too. It works. It's fantastic. I'm getting better results than 27B now.

To add: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RC... IQ3_XXS is within a point of the fully unquantized model and IQ3_S actually beats the unquantized model on many tasks! You lose absolutely nothing. It is quantization magic :)

No magic needed, the quantization found the expert which deals in java and deleted it, so the model overall became better.

lol. JDQ -- Java Delete Quantization it's the new thing!