Evals on actual research workflows is the right direction, most agent benches are toy tasks.

[deleted]

[dead]