It's not fully solved, there are still errors but its improved a lot the past few years especially compared to Youtube. The best approximation is the Whisper models atm as far as I have seen. There is a write up here on hugging face which talks more about the models and the benchmarks against it.

https://huggingface.co/openai/whisper-large-v3