I have been playing with a lot of $20 devices with no keyboards or screens. I wanted to use speech to interact with them, but without it having to run through a cloud service or use a fixed command vocabulary.

I worked with Claude to implement high performance 4bit int kernels for the ESP32S3 and P4 and quantize some larger STT and TTS models. Surprisingly, it's quite usable on these tiny boards.

Also includes a finetuned model that can better cope with my diesel drones, hacking coughs, and mumbling. Works with Micropython (RAM scarcity issues) and AtomVM (better).

I was hoping this could be a hardware STT-to-HID, but, even after looking at the Git ReadMe, I still don't understand how is the text supposed to exit the device?

It's a library for C / Micropython / AtomVM, so you could send the result over a lora radio, ESPNOW, WAN, match it in a switch statement to activate a relay, draw on the LCD.. that's up to you.

Many ESP32S3 dev boards have a microphone and speaker header available conveniently on the board, so I have been using that to respond via Babytalk's text to speech capability.

I'll improve the README to make it a little more clear.

Thank you for the explaination. So it is a library and would require someone to adjust it to their use case. That's what got me confused: wide open use case, but nothing tangible. I mean, if an ESP device is doing STT+TTS, then what am I supposed to be talking to? But if it is an ESP device that can do ANY->TTS and/or STT->ANY with some code customization needed, then I think I get it.

I wrote the README from the perspective of being understandable by the lay person, and I think I left out some vital clues about this being a software building block for making your own ESP32 projects more intelligent. I pushed a README update to make it 0.1% more clear.

As for what to talk to.. or the ultimate purpose lol.. you could do simple device control scenarios ("lights off"), do hands-free sensor readings ("temperature at 95 degrees"), change wifi settings via voice ("switch access points"), etc.

I use it to control an agent-powered diverse device sensor network in my home and my truck. Much of the capability needs Internet access, but having on-device speech to text and text to speech means I can still do some stuff when I didn't bring the Starlink with me. On-device STT opens up a lot of low bandwidth (lora) opportunities too.

Thanks. I will ask an LLM to investigate if this could be implemented in a STT-to-HID standalone offline device. If I reach any semblance of success, I will let you know so the projects can link to each other.

- need a real time original voice to clone voice open source offline library

- need it for recording gaming videos while talking into mic with my voice but output is clone voice

[dead]