Hand-crafted, domain-specific tools massively outperform general purpose ones.

Shell execution and raw DOM access are great for a backstop, but you can go so much further with just a little bit of translation and delegation around the environment.

I think browser automation is probably the most apt scenario. Often a human who understands how a page is meant to be perceived can transform a megabyte of raw web content down to a few hundred bytes of plaintext without any reduction in fidelity. This can be achieved using deterministic code that is guaranteed to provide a perfect transform every time.

The performance difference between raw DOM access and curated plaintext is like a step function. With raw access you get maybe 10-15 steps into a complex workflow before the wheels pop off. With curated access I've seen it go 100+ screens without issues.

Simply managing the token bloat is probably the most important objective here. If that's all you focus on it will probably go really well.

Yes exactly. It's almost like there's this expectation that if the LLM GPU-Maxxes enough it can do anything but like... why. There's no reason to wade through a million HTML tags if you can slice out the exact div you want to check