I really wish GLM models had vision capabilities. I've worked around that in the past to use a vision MCP in my harness that GLM can call. It is not the same, but it allows the model to query images.

Well, now one of them does!

That's wonderful. I was going off an older version of the Artificial Analysis page for GLM-5.3-Flash https://artificialanalysis.ai/models/glm-5-3-flash. The page is updated now to show that it does support multi-modal image inputs.