If the performances are comparable, and there is no evidence it's not.
in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47
That is a massive cost reduction.
Refs: https://www.alibabacloud.com/help/en/model-studio/model-pric... https://runware.ai/gemini-omni
You can't just look at the per token cost, but how many tokens it takes on average to do a task. The difference can be massive.
Also cache write/read cost + cache efficiency.
True, but it would have to be more than massive (order(s) of magnitude) to offset that gap.
We notice with frontier models like Astra and Fable that one might use a lot less tokens than the other to complete the task thereby being the better deal in spite of the far higher token cost.