TL;DR: If you are interested in TTS, you should explore alternatives
I tried to use it...
Its python venv has grown to 6 GBytes in size. The demo sentence
> "This high quality TTS model works without a GPU"
works, it takes 3s to render the audio. Audio sounds like a voice in a tin can.
I tried to have a news article read aloud and failed with
> [E:onnxruntime:, sequential_executor.cc:572 ExecuteKernel] Non-zero status code returned while running Expand node. Name:'/bert/Expand' > Status Message: invalid expand shape
If you are interested in TTS, you should explore alternatives