I've got a standard check-list of tests to run before any deep analysis on faulty circuit behaviour:
1. Check the power rails are correct and stable (at the target devices, not the PSU source)
2. Check the resets are correct (polarity, level, sequencing) and reaching where they're needed
3. Check the clock is toggling cleanly (no jitter, has clean monotonic waveform)
4. General signal integrity and setup/hold of signals that are related to the issue
That catches a lot of basic issues. If those are all clean, then you go deeper.
To the article author's suspicion of crystals, I have seen crystal oscillators fail (stopping toggling) as ambient temperature ramps up and down; that's a nasty one to catch and prove, but it can happen. Changing vendor was the only solution there.