Yesterday night I was doing a project with QwenTTS 1.7B.
After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).
I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.
So yeah the cat is out of the bag for sure.
Yeah I was doing that at the beginning of this year with voice samples locally from hollywood-stars with Qwen3-TTS.
It took like 1min - capture something from a youtube or video and put in your own text. It worked also really good for a german test.
Made a voice message for my wife from one of our favorite actors, telling here how nice it would be to make some breakfast :D
Any chance you could share a bit more detail? I’d love to try this myself but could use some proven structure / approach.
Same, I'd love a link.
Its so crazy to me how prevalent bad AI voices are, when local models can do such good AI voices
Arguably, no AI voice should sound like a human voice:
https://youtu.be/M-IVVJkZnuo?t=236
Very few people explore the options they have and tend to stick with the first thing that works.