Have you compared this to using GPT-6.1 Sol instead of GPT 6 Astra + Deepseek? From my test, 6.1 Sol is a lot more token efficient than 6 Sol while being similar to Astra in performance, and I don't really find 6 Astra to be significantly better than 6/6.1 Sol for general coding as I feel 6 Astra is only noticeably better at spatial reasoning/vision compared to 6 Sol, and 6.1 Sol really closed the gap on that front.

6.1 Sol though is horribly slow. The time cost alone offloading to ds flash is probably worth a look.

I don't really mind the speed, I like watching Codex work most of the time, slower work means I have time to do corrective nudging for when the initial prompt was unclear.

It mostly feels horribly slow if you leave reasoning effort at max for Sol 6.1. If you dial it to normal/high/xhigh, Sol 6.1 is only a couple of minutes behind Astra, performs almost as well on terminal bench v4 tasks, and is about 4 to 5x cheaper.

No its like 30 t/s… its really slow.