I can’t speak to the veracity of the benchmarks but it appears their methodology is sound. Nanobot has 47k stars, fwiw https://github.com/HKUDS/nanobot

It has been more of an OpenClaw or Hermes alternative than a coding agent like OpenCode or Pi, so it’s likely to do well given less context bloat.