There are, broadly, two kinds of audio recordings that linguists want to transcribe. One is native speakers telling traditional stories, where they're speaking naturally and taking the natural shortcuts (such as "wanna" and "gonna" in English). The other is native speakers reading words (or short example sentences) very carefully and distinctly, so that the linguist can listen to the recording over and over to learn how to pronounce the word right. In those recording, they'll say "want to" and "going to" rather than "wanna" and "gonna".

Thanks for the pointer; I'll check out that model and see if it handles the "slowly and carefully" type of recording better than the "natural speaking" type. (And depending on what kinds of errors the model makes, even the recordings where it makes errors can prove useful: for example, a linguist studying regional variations in speech would want the model to produce the IPA for "gonna" rather than "going to").