Yeah that is a very good point, it’s something I thought about a lot when positioning the app. Essentially there are 2 mitigations to this, one is that it’s not really ideal for beginners and mainly targeting intermediate learners who can recognise or at least question errors to a certain extent but don’t want to create the entire subtitle file themselves. There is an inline editor for them to fix the errors.

The other thing is the speech recognition models are very sensitive to the audio. If the audio quality is good then the model would do a good job, I’m not sure how much experience you have with Whisper Large but it is very capable on normal speech. The issues arise when there are many competing sounds overriding each other