Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models.

What's your exact use case?

Sound effects are completely different to voice, which those models are trained to output.