- i do game recordings

- my actual voice is terrible and i have a lot of noise in vicinity to sit and process it

- what i want is to be able to talk into the mic and an AI voice is generated in real time with a very large degree of accuracy

- bonus if it works with OBS as a plugin or something

That's interesting -

I have been playing with Whisper speech-to-text recently (both CLI, and also . Depending on the model & computer performance, it seems like it can stream (listen and transcribe to text) audio. Maybe you can pipe that to a text-to-speech service (remote or local), and configure that as an audio input to OBS.

Are you live-streaming, or just recording video file to edit/publish later? If it's not live, that may relax some of the latency requirements.

Related links:

- https://github.com/openai/whisper (speech-to-text for terminal)

- https://buzzcaptions.com/ (graphical interface to Whisper)

Somewhat related, Descript mentions it does dubbing. I recently used this tool to edit a video, and I was impressed by how easy it was, and accurate too. It may not fit your realtime use-case, but might be worth exploring...

https://descript.com/

Good luck!

- i am curious about something else as well

- what exactly is the architecture of speech to speech?

- is it speech to text and text to speech or am I badly understating the actual mechanism

thank you for the links. i dont intend to live stream but to record