4.1 flash is very fast and capable. Token efficiency is not great so it fill up context window much faster compared to similarly capable models.
glm 5.3 flash is a tad slower but a bit more capable and way more token efficient.
Source: self hosted tested on rented GB200 node at 8bit.
Wow, I'm surprised you are saying GLM 5.3 Flash is more capable. Isn't is like half the price of 4.1 Flash?