Actually there are models for classifying sections in audio files as speech or music, for example Silero VAD or YAMNet. I am rather confident that any decent LLM can one-shot a script that downloads the model, feeds it all the podcast episodes to generate timestamps where music is playing, and uses ffmpeg to cut out those parts. For a nice touch maybe replace each music section with a few seconds of silence, add a few second fade-out to the part preceding the silence and a few second fade-in after.