One month coding with GLM 5.3 Flash

A developer shares their experience using the GLM 5.3 Flash model for coding tasks over the course of a month. The experiment highlights the challenges of model selection, cost management, and infrastructure availability in an agentic workflow.
Why it matters
It provides a realistic look at the operational costs and technical hurdles of integrating LLMs into professional software engineering pipelines.
Setting a challenge to spend the whole of September on only one efficient open model felt like a great idea at the time. Turns out not so much in practice. 2B tokens later, here’s how it went.
Here’s the tokens distribution according to AgentsView, one of our Agentic engineering recommendations to keep tabs on AI usage:
Zooming in on the models split specifically:
The goal was to spend the whole month on GLM 5.3 Flash pictured in teal. Here’s what went well:
The second half of the month didn’t go so well, with 1B tokens going to other models.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in