Hacker News

> Are there specific benchmarks that compare models vs themselves with and without scratchpads? High with:without ratios being reasonier models?