A new bot called Codex Astra has taken up StarCraft’s Brood War as its hobby, and it has emerged as the best AI in the game. A recent benchmark, known as Brood War Bench, matches AI models against one another in the classic real-time strategy title, and Codex has come out on top. The rest of the field cannot even reach a beginner’s level of play.
The benchmark, uploaded by the experiment’s author, reveals Codex as the clear winner over every other model, while Grok models have yet to prove they can play Brood War. The remaining systems lag behind, mostly occupying themselves with processing rather than action.
Codex’s Cheese Problem
The most persistent concept in Codex was disruption, and it took the form of a Probe moving across the map to target enemy workers or buildings. This tactic proved surprisingly effective, since the opposing agents frequently spent thirty to forty seconds contemplating how to respond to the probe while failing to act on anything else.
Codex’s systems were much weaker at steady production. It delayed tech, dribbled one or two basic units into defended bases, and threw workers into last stands. It also created distinct subagents for the economy, army production, and army control, which communicated little with each other. The army agent often sent each new unit straight into an attack without knowing what larger force the other agents were planning to build.
One frequent error among newcomers is to send units one by one rather than waiting for a larger group and a prepared attack timing. The person who worked on Codex said the game became much stronger at planning these moments, coordinating its subagents to act in concert when he helped guide its development.
The endurance held firm. Within G009, once the army and primary base were lost, Codex 5.6 Terra / medium carried its last Command Center and set it down in the far corner. It endured for another six minutes.
Grok’s Long Stretches of Reasoning
Grok 4.6 often produced long stretches of reasoning and very few command batches. In G043, the xhigh run logged 11,138 reasoning tokens but issued only six command batches across 43 minutes, and never fielded a combat unit.
Grok / xhigh made three Marines in G003, and Grok / medium made two Zealots in G002, yet neither reached the enemy base. These two cases look less like poor strategies than failures to keep observing the map and acting on what was seen.
Fable’s Ambition
I found myself cheering for Claude Fable in more than a few matches. Fable tended to focus on developing an economy and advancing up the tech tree rather than settling for the first unit it could muster. It appeared more engaged with actually playing the game than any of the other models.
The game reached a Lair, Spire, and Mutalisks and won in G007. Then, in G027, it added a Robotics Facility, Citadel of Adun, Observatory, and Templar Archives before the win. Ambition did not guarantee execution: in G036 Fable reached a Factory and Academy but Opus 5 overran it.
Neither Astra nor Fable managed to construct complicated armies, defend basic attacks, or play clear strategies. A brand new player engaging in photon rush would easily triumph in every single one of these encounters.
What the Benchmark Actually Measures
Every model was pitted against every other in a round-robin grid, with the contests held in parallel across Freestyle VMs. The system saved data from the game engine and the logs for both agents for each match.
Playing a homemade version of Brood War with friends revealed how far the models can go on their own. After the creator built this agent-only version, they played it with a couple friends who had only played a couple Starcraft games in their lives. They did surprisingly well, and when the creator asked why, they said they hadn’t done much — they had asked their agent to attack, and it had built a small army and done the full attack for them.
That is where the benchmark came from. The person who made it wanted to see how far the agents could go on their own.
Codex vs. Grok vs. Fable
| Model | Strength | Weakness |
|---|---|---|
| Codex Astra | Consistent leader | Sends units in one at a time |
| Grok | Long stretches of reasoning | Never fields a combat unit |
| Fable | Climbs the tech tree | Ambition without execution |
The Persistence of Codex
G009 showcased Codex’s ability to endure. After losing its army and main base, Codex 5.6 Terra / medium raised its final Command Center and moved it toward the opposite corner. It went on for another six minutes before meeting its end.
The notion that kept coming back in Codex was disruption. In Protoss matches, it frequently involved sending a Probe across the map to assault workers or buildings. The plan worked surprisingly well, since the opposing agents usually spent thirty or more seconds contemplating how to respond to that probe rather than taking any other action.
These systems performed poorly at sustained production. Codex delayed technology, releasing one or two basic units at a time into defended bases, and threw workers into last stands. It also produced separate subagents to manage the economy, army production, and army control, which communicated little with one another. The army agent often sent each new unit straight into an attack without knowing what larger army the other agents were planning to build.
A frequent error among newcomers is to send units one at a time, rather than waiting for a critical mass and a planned attack timing. When the creator helped direct Codex, it was much better at planning those moments and getting its subagents to work together.
What This Means
The benchmark is a snapshot of where these models stand in the strategy game domain. Codex Astra’s cheese-first approach shows a model that can disrupt with the best of them but struggles with sustained production and coordination. Grok’s long stretches of reasoning show a model that thinks deeply but rarely acts. Fable shows ambition but struggles with execution.
None of the models went beyond a beginner level, yet the creator’s own reaction was one of excitement: “This benchmark is nowhere near exhausted. There is much more for the agents to learn, and much more for the benchmark to ask them to do.”
The creator found themselves more excited than usual while watching the agents play. This is the most honest way to describe the whole experiment: the bots are terrible at Brood War, but their antics are genuinely entertaining to watch.
The Contrast Between Codex and Grok
The difference between Codex’s disruption and Grok’s heavy use of reasoning and the benchmark makes the latter feel less like a scientific experiment and more like a comedy act featuring bots.
These models have not reached beginner-level performance, but they are moving in that direction. Codex’s lasting power, Fable’s bold reach, and the vast amount of reasoning shown by Grok all prove that these systems are not standing still. Each one is merely standing still in a distinct way.
The results have arrived, and the benchmark shows Codex Astra as the winner, beating every other model with consistency. Grok models have not yet reached the level needed to play Brood War. The creator is already anticipating the day when they do.
Source material: “Brood War Bench,” swerdlow.dev.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

