Midterms 2026See who we think should earn your vote, based on our standardsThe guide →
WRITTEN IN PLAIN AMERICAN ENGLISH.
CLAY TRIBUNE.
Advertisement

Team Spent September Trying To Code On One Model — And Failed

A team spent September using one efficient model for all work, failing to stay within budget and learning hard lessons.

By mitch·5 min read
A workstation glows with model charts and token counters, symbolizing a team's failed budget experiment.

A small team spent September trying to use one model for everything they did. They wanted to prove they could do engineering on a single efficient open model, month after month. They failed. They used twice as many tokens as they planned.

The story, titled “One Month Coding With GLM 5.3 Flash,” is a cautionary tale about model choice, infrastructure, and the hidden costs of experimentation.

The Plan

The goal was simple: spend the whole of September on GLM 5.3 Flash, an efficient open model with a large context window and vision support. The team wanted to see if they could keep their work lean while keeping up with what providers were releasing. They set a budget, tracked their usage through AgentsView, one of their Agentic engineering recommendations, and watched the numbers closely.

Advertisement

“It turns out not so much in practice.”

That is the opening line of the post, and it sums up the whole month. The plan sounded good on paper. In execution, it did not hold.

What Went Well

The first half of the month went smoothly. The author stayed within budget, spending roughly $68, which works out to about 4kWh of energy use and 365 grams of carbon emissions. GLM 5.3 Flash performed well on Wagtail itself, on sites built with it, on UI tasks, AI R&D, and documentation writing. It was polyvalent, handling extended coding tasks, screenshots, and QA visually. The model’s availability across a wide range of providers meant the author saw the benefits of healthy competition.

The author’s assessment of the model’s strengths is worth repeating:

  • A 1M context window, so any extended coding task is possible.
  • Vision support, so it can build off screenshots or QA visually.
  • Available across a wide range of providers, so you see the benefits of healthy competition.

The Unexpected Hurdles

The second half of the month did not go as planned. The author spent 1B tokens on other models. There were three main reasons for the overrun.

Vibe Coding

The author’s experimental Wagtail MCP server is a vibe-coded prototype. Vibe coding is not what they normally aspire to, but for a prototype it worked. The problem was that they chose the wrong model for the prototype. They spent 450M tokens, or $150, or 5kWh of energy use, almost overnight.

The MCP server itself works well, and they now have a great demo of the capabilities. It is not for nothing. But the cost was real. They could have achieved similar results for most likely 5x less cost with not that much more effort. The lesson is straightforward: be careful with model selection and with agentic patterns. Budget for this, and be more careful.

Infrastructure Woes

The author has written extensively about comparing inference providers. Their choices work most of the time, but they are very popular. They do not have the same capacity as the big labs that hoard all the GPUs. The author noted degradation with the performance of GLM 5.3 Flash in particular, most likely because of it being so high up the Pareto frontier of relevant models for their work.

This meant having to switch to other similar models: DeepSeek V4.1 Flash and Qwen 3.8 Flash. Switching is simple, but it was still unexpected. The author had to move away from their target model because the infrastructure could not keep up with demand.

Experimentation Costs

Beyond day-to-day engineering, the author felt it was essential to keep experimenting with a wide range of models. They needed data across a wide range of models as they started to benchmark performance on Wagtail tasks. The benchmark is a sneak peek, and it shows the value of concrete data.

It is much easier to guide people toward leaner options with this kind of information. The author is also working on a new CLI prototype intended to work well with agents. Having that data helps make those options more viable.

The Cost of Failure

Only 50% of the month’s usage was on the target model. That is a technical failure by any measure. The author spent 1B out of 2B tokens, and about 35 kWh of energy use instead of 10. The overrun was significant, and the author knows it.

But the failure produced something valuable: lessons learned. The author reflected on what went wrong and what they need to do differently for October. The reflection is the core of the post, and it is worth reading carefully.

Lessons From September

The author identified four key areas for improvement:

  1. Constant, local usage measurement and reporting. Look not just at tokens but also at energy use and spend, and ideally at how well all of this leads to concrete positive outcomes.
  2. Budgeting for experimentation, not just day-to-day tasks. Make more concerted decisions about which prototypes are worth building and how.
  3. Better prompt selection and multi-agent techniques. Use orchestrator, scout, implementer, and reviewer agents. Set bounded goals. Not rocket science but certainly one more thing to learn.
  4. Keep pushing for more efficient techniques and models. The Jev-style decision diffusion models look very promising if they can run so efficiently. The latest flagship models also look like a step in the right direction on that front.

The author wants to keep the challenge going, but only for the 50+% of normal day-to-day production, not for R&D. The R&D work requires a wider range of models to benchmark and test. The production work can stick to one or two flash-tier cheap models.

What To Try Next

The author’s goal for October is simple: the majority of AI inference work should be done with such efficient models, measured in cost or energy use rather than meaningless tokens. That is the target they are setting for themselves.

They are also inviting others to join in. The post closes with an invitation to come say hi at Wagtail Space 2026 in November to hear how that all pans out.

The post is a reminder that planning and measurement are not enough. You also have to watch what happens, and you have to be ready to adjust when the numbers stop matching the plan. The author did that, and they wrote about it honestly.

See the video the story is built around at wagtail.org.

The Notebook

Get the Notebook.

The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

We send one note to confirm. Every issue has a one-click way out.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

As an Amazon Associate, Clay Tribune earns from qualifying purchases.