Yesterday night I was doing a project with QwenTTS 1.7B.

After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).

I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.

So yeah the cat is out of the bag for sure.

Yeah I was doing that at the beginning of this year with voice samples locally from hollywood-stars with Qwen3-TTS.

It took like 1min - capture something from a youtube or video and put in your own text. It worked also really good for a german test.

Made a voice message for my wife from one of our favorite actors, telling here how nice it would be to make some breakfast :D

Any chance you could share a bit more detail? I’d love to try this myself but could use some proven structure / approach.

Same, I'd love a link.

Its so crazy to me how prevalent bad AI voices are, when local models can do such good AI voices

Arguably, no AI voice should sound like a human voice:

https://youtu.be/M-IVVJkZnuo?t=236

Very few people explore the options they have and tend to stick with the first thing that works.

[deleted]