Deepseek's DSpark does dynamically adjust speculated token count per user/completion.
https://arxiv.org/abs/2607.05147
But they do so to maximise total throughput, I don't think there's reason to do that for batch=1.
Deepseek's DSpark does dynamically adjust speculated token count per user/completion.
https://arxiv.org/abs/2607.05147
But they do so to maximise total throughput, I don't think there's reason to do that for batch=1.