Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models.
What's your exact use case?
Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models.
What's your exact use case?
Sound effects are completely different to voice, which those models are trained to output.