Update: tested it out myself on their Max plan, on some parallel agentic sessions.
Currently 20% of my 5 hour limit and 4% of my weekly limit.
Total: 58.46M
GLM-5.3 Cached: 56.91M
GLM-5.3 Uncached: 1.23M
GLM-5.3 Output: 315.18K
Cache hit rate: 97.9%
Extrapolating from that (inaccurate for now but oh well): Full 5-hour Full weekly
Total 292.3M 1.461B
Cached 284.6M 1.423B
Uncached 6.15M 30.75M
Output 1.576M 7.88M
All of the work was off-peak I think, using OpenCode not ZCode in these examples.Their own estimates are quite different, probably due to their conservative caching estimates vs what I normally get on longer form work: https://docs.z.ai/devpack/overview#estimated-token-allowance