Because byte-for-byte identical output is far easier to measure than trying to test for functional equivalence.

This is also partly a preservation activity so (as best we can create it) identical code generating identical output is a big part of the point.