Most of My Antifragile Production Wins Were Accidents, and I Only Noticed by Looking Back
Reactive hardening is not the same as seeking stress
Antifragile production means the management system gets stronger when the plan breaks
The difference between robust guardrails and genuine stress-seeking mechanisms
Why most of my hardening is reactive, and why that is the cheaper half
The one move that turns a mistake into something that outlives the ticket
The criticism of AI in production usually points at the model. It makes things up, it drifts, it cannot be trusted with anything that matters. All fair. But it aims at the wrong target. The fragile thing in game production was never the model. It was the plan.
I have spent my career in game production watching producers, myself included, pour enormous effort into artefacts that shatter on contact with reality. The Gantt chart is a prediction that is wrong the moment it is published. The risk register is a document that goes stale the day after the offsite. The milestone plan assumes a stability that games have never once provided. We maintain these things by hand, heroically, and they break anyway. Only 30% of software estimates come in accurate, and developer accuracy does not improve with practice. A McKinsey study with Oxford of more than 5,400 IT projects found the average large one runs 45% over budget and delivers 56% less value than promised, with 17% going so wrong they threaten the company itself. This is the normal state of things, not the exception. The management layer, the thing whose entire job is to protect against chaos, has always been the most fragile part of the operation.
If this is useful, there's more where it came from. Multiple free posts a week on game production, applied AI, leadership and how the sausage actually gets made.
So when I started running production work agentically, the interesting question was never whether the agents would make mistakes. Of course they would. The question was whether I could build a management system that got better every time the plan was wrong, instead of one that just absorbed the hit and carried on.
That is the whole idea behind antifragility, and it is worth being precise about it because the word gets thrown around loosely. Nassim Taleb‘s spectrum runs fragile, robust, antifragile. Fragile breaks under stress. Robust survives stress unchanged. Antifragile gets better because of it. His line is the one to hold onto: the resilient resists shocks and stays the same, the antifragile gets better and better. Most people stop at robust and think they have arrived. They have not. Robust is a system that shrugs off a dip in throughput. Antifragile is a system where the dip changes how you plan the next stretch of work.
I want to be clear about what I am applying this to. Not the game. Not the art, the code, the levels, the assets. The producer’s own kit: the schedules, the trackers, the standing rules, the routing, the memory of how the work actually goes. That layer. The one that has spent decades pretending to be robust while being fragile underneath.
When I looked back over my own agentic setups through this lens, I found something uncomfortable and useful. The antifragility was real, but it split into three tiers, and the tiers were not equally earned.
{{button:subscribe}}
The first tier is the genuine article: mechanisms that go looking for stress before it arrives. On one project I stood up a dedicated red-team role whose entire job was to manufacture pressure on assumptions and designs before they shipped. Not review them politely. Attack them. That is antifragile by design, because the system actively seeks out volatility rather than waiting to be hit by it. Alongside it, an open-questions tracker that treats every unresolved uncertainty as a first-class artefact, surfaced continuously, never quietly assumed away. The spec gets sturdier each time a gap is logged rather than papered over. And a skill-gap mechanism that detects missing capability under real task pressure and grows the team in response, so capacity increases because of the strain, not despite it.
Those three have something in common. They do not wait for a failure to teach them. They generate the stress themselves, on purpose, while it is still cheap.
The second tier is where it gets honest. This is reactive antifragility, and when it works it is the most satisfying thing in the whole operation. The clearest example I have is an export pipeline that failed, repeatedly, in every way a pipeline can. Full-page captures when it wanted a frame. Text that would not scale. Frames stacking on each other. Stray markers leaking through. Each failure got converted into a permanent rule, written into a lessons-learned list baked into the process doc. That pipeline is measurably stronger now, and it is stronger precisely because it broke early and often. Nobody designed it to be good. It became good by being repeatedly wrong in public and refusing to forget any single instance.
The pipeline did not survive its failures. It fed on them. Every break left the system harder to break the same way twice.
That is closer to true antifragility than any amount of careful upfront planning, because the failures improved the system rather than merely being patched.
The third tier is the one I have to be careful about, because it is the one most likely to flatter me. This is reactive robustness wearing the costume of antifragility. Consolidating onto one tool after a parallel workflow went redundant. A confirmation rule for parent pages. An auto-numbering convention. A carve-out that exempts certain items from a process that did not suit them. Every one of these is a response to a specific incident, and every one of them closes off a failure mode so it cannot recur. They are good. They are also just guardrails. They make the system harder to break again, but they do not make it get better from the next, different stress. A guardrail that stops you repeating one mistake does nothing for the mistake you have not made yet.
Most of my antifragility lives in that third tier, and a good chunk in the second. The genuinely proactive, stress-seeking mechanisms, the red-team role, the capability-gap process, are the rarest and the least used of the lot. I have built far more documentation-hardening than stress-seeking, because documentation-hardening is cheap. Something breaks, you write a rule, the rule compounds, and it feels like a system that learns. It looks like antifragility from a distance. But reacting well to mistakes you have already made is not the same as going out to find the ones waiting for you. The cheap half is doing most of the work, and I only noticed by counting.
Thirty-plus years of production work went into this. Most of it didn’t make the post. New posts every week. Free to read, free to subscribe.
The antifragile play in agentic production is not “avoid mistakes.” It is turn each mistake into a rule that outlives the ticket it came from.
The mistake is not the enemy. The mistake that teaches nothing is the enemy. A mistake you convert into a standing rule has paid for itself.
That sounds small until you watch what it does across a project. A sub-task format that went wrong twice on the same ticket does not just get fixed on that ticket. The correction becomes a standing rule, and the whole class of error stops happening. A live change that went out before I had actually approved it becomes a review-before-verify rule: drafts get read as text first, real action only follows an explicit go-ahead. A labelling slip becomes a labelling convention. A one-off mapping gap gets closed by building the mapping document itself, which then prevents the drift for good. None of these fixes stays local. Each one outlives the incident it came from and becomes part of how the system behaves.
This is why, when I set these projects up, the memory has a dedicated category for feedback, kept separate from one-off project notes. That separation is the entire point. If the lesson lives only in a ticket comment, it dies with the ticket. If it lives only in my head, it walks out of the door the day I do, and every producer reading this knows exactly how much institutional knowledge has left their studio that way. The rule has to live in the system. That is the difference between a producer who is good and a production system that is good, and only one of those survives your holiday.
There is a boundary here, and it matters more than the technique, because it is the thing that separates antifragile from merely chaotic. Antifragility only pays when two conditions hold. The downside of each experiment has to be capped, and the signals you learn from have to be independent.
The capped-downside part is where good intentions go to die. An agentic layer that re-forecasts the ship date every time throughput dips for a week is not antifragile. It is a sophisticated panic machine. It reacts to noise as if it were signal, and it will exhaust the team faster than any fragile Gantt chart ever did. Antifragility is not “react to everything.” It is small, controlled exposure to disorder with hard limits on how far any single event can move the plan. The red-team role works because its blast radius is contained. If a failed experiment could corrupt the main line or eat three weeks, you do not have an antifragile system. You have an expensive fragile one wearing a Taleb t-shirt.
The independence part is subtler and easier to get wrong. If you learn a lesson from stress that turns out to be a fluke, you have hardened your system around a coincidence, and that is worse than not learning because it feels like progress. The same trap catches the agents themselves. Check one agent’s output against another instance of the same model and you get agreement, not verification. Agreement is cheap. Real verification checks the claim against something outside the model: the actual board, the real telemetry, a human who looked. Independent signals harden the system. Signals that only echo each other harden it around nothing.
None of this is theoretical. Across the wider economy 79% of organisations have deployed AI agents at some level but only 2% at full scale, and in games 36% of professionals now use generative AI while 52% hold a negative view of it. Some of that gap is the fragile version getting deployed: agents let loose with no capped downside and no independent check, doing confident damage at speed. The tool is not the problem. The absence of an antifragile frame around it is.
Point this question at your own setup, the same one I pointed at mine. Ask your project to look back over the last few weeks and find every place where your plan, your estimate, or your assumption turned out to be wrong. Then, for each one, ask whether that miss changed how the system plans now, or whether you just absorbed the hit and carried on with the same assumptions.
If the answer is a recovery story, you survived, and you are robust. If the answer is “the forecast range came in well off the actual because the data feeding it was messy, so I now clean that data before it drives the next forecast,” or “this dependency slipped twice, so I flagged it as a structural risk before you asked,” then the volatility fed back into the method, and you have built something antifragile. And if the answer is vague, all narrative and no changed behaviour, the learning has been living in your head the whole time, and that is the most common place for a production system to be secretly fragile.
The craft was never about writing a better plan. Nobody has ever written a plan a game did not eventually break. The craft is building a system that gets better each time the plan breaks. Agents finally make that cheap enough to do on purpose, and the only real work left is being honest about which tier you are actually on.
If this changed how you think about even one thing, the next post might too.




