This all rests on the assumption that design errors are independent beween LLM-generated programs for the same specification. That's not true for humans (Knight and Leveson, 1986) and I highly doubt it's any more true for robots.

(On the other hand I just found out about Ron, Baudry, Monperrus, 2026, which seems to say "sure, problems are correlated, but it could still be useful.")