Task I want but have been too lazy - I have a favorite podcast, that I often fall asleep to. Certain parts with music etc. wake me up. I'd like to chop those parts out as automatically as possible. Can I do that here by e.g. giving some example edits to a few files, and asking to remove similar stuff from other files?

Sounds like something a better model armed with ffmpeg would already be able to do. Run an analysis on loudness along the track, detect when speech begins/ends (with whisper) around loud segments, and compress or cut those parts out

Actually there are models for classifying sections in audio files as speech or music, for example Silero VAD or YAMNet. I am rather confident that any decent LLM can one-shot a script that downloads the model, feeds it all the podcast episodes to generate timestamps where music is playing, and uses ffmpeg to cut out those parts. For a nice touch maybe replace each music section with a few seconds of silence, add a few second fade-out to the part preceding the silence and a few second fade-in after.

an interesting use case... the analysis could detect rhythmic or not and then edit the rhythmic parts away. i will look into, added to my backlog. thanks for your feedback!

Cool! As the other comments suggest, I could think of plenty of ways you could do this, probably much more efficiently than having a GUI, once you figured it out. But I would still like something that felt like less software engineering or data science and more "this is how I would do it manually in Audacity, do this for me automatically on a batch of files"