This is really cool work! I'm curious like what do you see as the biggest lever for speeding up TTS models or from a technical perspective that this was a promising direction in the first place to push on. If I were to guess, some distillation but I'm certain there are probably TTS model aware architectural changes that just make inference wayyyy faster?

thanks for the interest! we have a blog post on exactly how we did it: https://narilabs.com/blog/qwen3-tts-speed-cost-frontier/