Update: tested it out myself on their Max plan, on some parallel agentic sessions.

Currently 20% of my 5 hour limit and 4% of my weekly limit.

  Total: 58.46M
  GLM-5.3 Cached: 56.91M
  GLM-5.3 Uncached: 1.23M
  GLM-5.3 Output: 315.18K
  Cache hit rate: 97.9%
Extrapolating from that (inaccurate for now but oh well):

            Full 5-hour  Full weekly
  Total     292.3M       1.461B
  Cached    284.6M       1.423B
  Uncached  6.15M        30.75M
  Output    1.576M       7.88M
All of the work was off-peak I think, using OpenCode not ZCode in these examples.

Their own estimates are quite different, probably due to their conservative caching estimates vs what I normally get on longer form work: https://docs.z.ai/devpack/overview#estimated-token-allowance