I Have Never Seen a Game Fail Because the Artists Were Too Slow
AI is optimising the cheapest part of production
Rework consumes 30 to 50 per cent of software development effort, and reworked code costs around 2.5 times first pass development
Skull and Bones burned eight years and a reported 200 million dollars on unstable direction rather than slow execution
Faster generation increases the volume of work in flight when a pivot lands, which increases the write off
Producers should measure rework rate, decision latency and decision durability before measuring tool speed
Every studio I have worked in could tell you the cost of an artist day. None of them could tell me the cost of a decision that got reversed six months later.
That gap explains most of what is currently happening with AI in game production.
Luke Dicken spent a decade at Zynga and finished up as Head of AI at Take-Two before his entire team was laid off earlier this year. He gave a long interview to 80 Level last week, and one passage has been rattling around my head since. Asked whether AI is a realistic answer to ballooning AAA budgets, he turns the question round and asks where the ballooning actually happened. How many projects went through repeated resets. How much of the money was spent building the thing that eventually shipped. Then he lands it: “No amount of automated asset generation is going to fix a broken steering wheel.”
I would push it a step further. The steering wheel is the one component nobody has bothered to instrument.
If this is useful, there's more where it came from. Multiple free posts a week on game production, applied AI, leadership and how the sausage actually gets made.
The money did not go where the tools are pointed
Software has numbers on this, and they are ugly. ScopeMaster’s summary of the research puts rework at 30 to 50 per cent of total effort on a typical project, and notes that reworking existing code runs at roughly 2.5 times the cost of writing it the first time. The NIST work from 2002 is older and blunter, putting around 80 per cent of development cost into finding and fixing defects against 20 per cent on preventing them.
Games are worse than the software baseline, because we carry content on top of code. When a design direction changes, the code churn is the small part. The animation set built against the old locomotion model goes. The level blockouts sized for the old encounter pacing go. The VFX authored against a colour language that has since been rewritten goes. The tools team’s export pipeline, tuned for a content shape that no longer exists, goes.
Which brings me to the exhibit I keep coming back to.
Kotaku’s investigation into Skull and Bones drew on more than twenty current and former Ubisoft developers. The core premise of the game changed at least four times. Three creative directors passed through, each rewriting the design in their own direction. Headcount grew from roughly 100 people in 2015 to around 400 by 2019, while basic questions about what the game was remained open. One developer’s summary of the experience: “Instead of adding layers of value we kept running around in a loop.” By 2021 the reported spend was 120 million dollars. By launch, Insider Gaming’s reporting put it nearer 200 million.
Nowhere in eight years of reporting does anyone suggest the problem was that the artists were too slow.
Four hundred people executing against a brief that keeps moving is a production failure with an enormous payroll attached. Nothing about it gets cheaper when the payroll works faster.
Skull and Bones is the extreme case, which is why it makes a useful exhibit. The ordinary version happens everywhere and gets no coverage at all. Embracer quietly cancelled 29 unannounced games in six months between July and December 2023. Nobody wrote a twenty-source investigation into those. Each one had a team, a set of decisions, and a moment when someone at the top of the chain changed their mind.
The part I have avoided writing
I have wanted to write this post for years and kept finding reasons to write something else, because I intend to keep working in this business. That hesitation is worth sitting with for a second, because it is the whole problem in miniature. The decision layer goes unmeasured because the people best placed to measure it report to it.
So, carefully, and without naming anyone.
For most of my career I have been the monkey in the middle. The gyration comes down from the C-suite as a sentence in a meeting. It reaches the team as two quarters of work. My job was to stand between those two facts and translate one into the other, usually while being asked why the date had not moved.
It is very easy, in a room with good coffee and no build running, to say that we will just make this one small change. It is a rock thrown into a pond by somebody who turns round and walks off before the ripples reach the bank.
The gyration comes down from the C-suite as a sentence in a meeting. It reaches the team as two quarters of work.
I have watched a single-player game with a clean spine and a real point of view get opened up for multiplayer because a competitor’s live service posted a good quarter. I have watched a multiplayer game that worked get committed to another platform, and then watched that platform’s constraints crawl back up the pipe and into the core design, one compromise at a time. I have built features that made no sense inside the game they were in, because a platform holder asked and publishing agreed, because the relationship was worth more than the build. Shit rolls downhill, and production is most of the hill.
The bit that does not show up anywhere is what it does to people. I have hired someone, told them the plan, and then told them a different plan, and then a third. There is a number of times you can do that to somebody before they stop believing anything you say, and it is lower than you would like. The strongest people go first, because they have options. What is left is a team that has learned to hedge and a producer who has learned to promise less.
It followed me home, too. Not in a dramatic way. In the ordinary way, where you spend a weekend rebuilding a plan you already know is provisional, and you do it again the following month, and the tiredness that comes from defending decisions you disagree with turns out to be a different and more corrosive kind of tiredness than the sort that comes from hard work.
None of that appears in a post-mortem. Post-mortems are written by the people who survived the project and read by the people who caused it.
Thirty-plus years of production work went into this. Most of it didn’t make the post. New posts every week. Free to read, free to subscribe.
What a reset actually costs
The naive model says a discarded feature costs whatever it cost to build. Producers who have lived through a few resets know that number is the smallest line item on the bill.
Start with the build cost, then add the integration work: the systems that were extended to accommodate it, the save format that was versioned around it, the UI that was restructured to surface it. Add the dependent work that other disciplines committed because the feature was on the roadmap. Add the review cycles, the pipeline changes, the tooling written to support content that will never ship. Add the opportunity cost of whatever the team did not build instead. Then add the quiet tax at the end, where every subsequent estimate from that team comes back padded, because they have learned that commitment is provisional.
That last one compounds across a studio and never appears on a budget line.
The important property of all this is that it scales with how deeply the work was integrated, and integration depth is a function of how long the wrong direction ran before someone corrected it. Which means the cost of a bad decision is set by decision latency. A direction reversed in week three costs a prototype. The same reversal in month fourteen costs a milestone, a pipeline, and usually a few senior people who decide they would rather work somewhere that finishes things.
The most expensive rework lands at beta
Beta is often the first time the executive layer sees the game the way a player will see it, start to finish, in something close to final form. Everything before that was a slice, or a build with a jump-to-level menu and a producer narrating over the top.
Two things then happen with grim reliability.
The first is a quality uplift demand, and the cruel part is that nobody did anything wrong to cause it. Content built in the first year was approved on its merits at the time. The team was still learning the toolchain, the shader library was half written, the pipeline had not settled. By beta, the last areas built look considerably better than the first, because the team got good, which is what teams are supposed to do. Then somebody plays it end to end and asks, fairly, why the opening hour looks worse than the closing hour. The answer is that it was made first. That is not an answer anyone accepts.
So the studio goes back and redoes its earliest content at the exact point in the schedule where every hour is contended and every system that content touches has been locked. On most projects I have worked on, that single pass is the largest rework bill of the entire production, and it was set in motion by an approval that was correct when it was given.
Nobody made a mistake. The early work was approved correctly, the team got better, and the bill arrives anyway.
There are two honest ways to handle it. Treat the quality bar as a moving target, budget an uplift pass from day one, and defend that budget when the schedule tightens. Or hold the bar fixed, approve early content against the standard the team can actually hit in year one, and accept that the opening will look like year one work. Studios that decide neither thing pay for both.
The second is the first-time user experience, and anyone who has shipped a mobile title carries this scar. FTUE has to be built before the game is finished, because it ships with the game, and it has to teach systems that are still moving underneath it. It is the most dependency-dense content in the build. Every economy tweak, every UI restructure, every rebalance of the early progression curve, every system that arrives late in the schedule reaches into it.
And it gets rebuilt. Sometimes because soft launch retention came back wrong, which is a legitimate reason to rebuild it. More often because the numbers are not landing the way somebody expected and FTUE is the most visible thing available to change, or because a deck went round with this quarter’s onboarding best practice and now the game needs a different one. Each rebuild is cheap to request and expensive to absorb, and it lands in the window where the team has the least capacity left to absorb anything.
Faster execution makes an unstable brief more expensive
Here is where I part company with the cheerful reading of production AI.
If direction were stable, halving execution cost would be a straightforward win. Direction is rarely stable. And in a pipeline with unstable direction, faster execution has a second-order effect that goes the wrong way.
I spent a decade arguing with studios about work in progress limits, usually to blank stares, occasionally to open hostility. The argument is simple. Everything committed and unfinished is inventory. Inventory carries risk. When the direction changes, everything in flight becomes a decision: finish it, bin it, or park it and pay to remember it later. Raise throughput without raising the rate at which direction settles and you raise the amount of inventory sitting in flight at the moment the pivot lands. There is more to unwind, more integration to unpick, more of the dependency graph poisoned.
Cheap generation does not make untangling cheap. Nothing about a faster asset pipeline reduces the cost of the decisions built on top of the asset.
This is the part that gets missed when a studio measures its AI deployment. The measurement is almost always taken at the task level, where the tool is doing its job perfectly well. Concept iterations per week went up. Time to first playable audio went down. Every one of those numbers can improve while the total cost of the project rises, because the numbers describe the cheap part of the system and say nothing about the expensive part.
The counter-argument I take seriously
There is a good rebuttal here and it deserves a fair hearing.
If execution becomes genuinely cheap, then being wrong becomes cheap, and a studio can afford to explore more directions before committing to one. Run six prototypes instead of two. Kill five of them in week four rather than month fourteen. That is a real benefit, and it points at the one place where I think production AI earns its keep right now: pre-production, where the artefact is meant to be thrown away and the throwing away is the point.
That benefit comes with a condition attached, and in my experience the condition almost never holds.
A prototype exists to answer a question. The answer is the deliverable. The code and the content are apparatus, and the right thing to do with apparatus is bin it and build the real thing properly, now that you know what the real thing is. Almost nobody does this. What happens instead is that the surviving prototype gets promoted. There is a date, the prototype runs, and proposing a from-scratch rebuild of something that already works is a conversation nobody wants to have with a publisher. So the thing built in six weeks to test one question becomes the skeleton of a four-year production, and every system added afterwards is shaped by decisions that were made to get an answer quickly.
Some of the hardest constraints I have worked under were put in place by people who were not making an architectural decision at all. They were answering a question fast, in the way you are supposed to answer a question fast. Somebody else then declined to throw the answer away.
Prototype code that survives contact with a ship date stops being a prototype and becomes an architecture. Nobody ever writes that decision down.
This is the part of cheap generation that makes me uneasy. Promotion pressure scales with how finished a prototype looks. A grey-box level with placeholder audio is obviously provisional and everyone treats it that way. A prototype assembled from generated assets and generated code reads as close to shippable, which makes the case for rebuilding it much harder to win in a room full of people looking at a screen. Dicken describes generative AI as automated scaffolding, useful for putting up a frame and incapable of building the house. The failure mode of scaffolding is somebody deciding it looks solid enough to live in.
There is a further wrinkle here that I have only started seeing in the last year, and I think it gets worse from where it currently sits.
Technical debt created by a human prototype has an author. The shortcuts go undocumented, but somebody holds the reason in their head, and three years later you can find that person and ask why the save system is shaped the way it is. The answer is usually short and usually makes sense once you hear it. That conversation is the only documentation the decision ever had, and it works often enough that our industry has quietly built its entire architecture practice on the assumption that it will be available.
Generated prototype code has no author in that sense. The person who prompted it can tell you what they asked for. They frequently cannot tell you why the result came out shaped the way it did, because the reasoning was never in anybody’s head to begin with. Regenerating it does not help either, since the same prompt does not reliably produce the same output. The archaeology a senior engineer would normally perform on a legacy system has nothing to dig into.
Human prototype debt is undocumented. Generated prototype debt is unauthored. Only one of those can be recovered by finding the right person and asking them a question.
Stack that on the two habits already in the building. Prototype code gets a lighter review because everyone agrees it is going to be thrown away. Generation makes it cheap to produce far more prototype surface than a small team could previously write by hand. Promote that into production and the foundation of a four-year project is code that nobody read closely, nobody can explain, and nobody can reproduce.
The fix is procedural and it belongs to production. If generated code is going to be promoted, the reasoning has to be captured at the moment of generation, while the person still remembers what they were trying to do, because that is the only moment it exists anywhere. Somebody has to own the question of what gets written down. That somebody is a producer.
Where the argument breaks completely is at the moment of integration. The expensive part of a reset was never the asset. It was the dependency graph and the accumulated decisions stacked on top of it. Generating the replacement asset faster does nothing for either. So the benefit is real in the phase where work is disposable by design, and evaporates in the phase where work is load-bearing. Studios are mostly deploying it in the second phase, because that is where the headcount is, and headcount is what the finance conversation is actually about.
Nobody has measured whether any of this worked
The uncomfortable backdrop to all of this is that the industry has almost no reliable evidence about execution-layer gains either.
METR’s randomised controlled trial, published in July 2025, put experienced open-source developers on real issues in their own repositories. Developers using AI tools took 19 per cent longer. The same developers, after finishing, estimated that AI had made them 20 per cent faster. Both numbers matter, and the second one matters more.
I want to be fair to METR, because they were fair to themselves. In February 2026 they revised their own design after finding that 30 to 50 per cent of invited developers declined to take part without AI access, which biased the original sample. Their larger follow-up cohort of 57 developers across 800-plus tasks showed a 4 per cent slowdown with a confidence interval spanning both directions. Their revised position is that we do not know. That is a more honest finding than the headline, and it is still devastating to anyone building a business case on self-reported speed.
Widen the lens and it holds. MIT’s Project NANDA reviewed over 300 publicly disclosed initiatives and found around 95 per cent of enterprise generative AI pilots produced no measurable impact on profit and loss, against 30 to 40 billion dollars of investment. Their diagnosis was integration and workflow fit rather than model quality.
Meanwhile the people doing the work have formed a view. In the GDC 2026 State of the Game Industry survey, 52 per cent of respondents said generative AI is having a negative effect on the industry, up from 30 per cent the year before and 18 per cent the year before that. Sentiment is not evidence. But when a workforce that uses a tool daily gets steadily more pessimistic about it across three consecutive surveys, that is data about something.
A studio that cannot measure its rework rate has no way of knowing whether its AI deployment helped, hurt, or did nothing at all. Most studios cannot measure their rework rate.
The three numbers I would want instead
If I were setting up measurement for a studio deploying AI tooling seriously, I would leave the task-level metrics to the tool vendors and instrument the decision layer, because that is where the money is going.
Rework rate
What proportion of completed work was subsequently discarded or substantially rebuilt, measured per milestone and per discipline. This is uncomfortable to publish, which is exactly why it works. It is also the only number that tells you whether faster execution made the project cheaper or merely made the write-offs larger.
Decision latency
Days from a question being raised to a decision being recorded. The meeting where it got discussed does not count. What counts is the point at which the team could act with confidence. Long latency means teams are building against assumptions, and assumptions are the raw material of rework.
Decision durability
How long a decision survives before it is reopened. A studio where the average creative direction lasts four months has a production problem that no toolchain will address. Track it and the pattern becomes visible within two milestones, usually pointing at a specific interface between two specific people.
None of these are hard to collect. All three are politically awkward, because they measure the layer of the organisation that commissions the measurement. That is the actual obstacle, and it has nothing to do with technology. I sat on this post for years for the same reason.
Dicken’s framing is that the bottleneck sits with the people telling the team what work to do. I would add the reason it stays there: it is the only part of the system where nobody keeps score. Every other discipline in a modern studio operates under measurement. Art has throughput and quality bars. Engineering has velocity, defect rates, build health. QA has coverage and escape rates. Production has burndowns and forecasts and Monte Carlo distributions that will tell you your ship date to a confidence interval.
The decision layer has status updates.
So we have arrived at a moment where the industry is spending heavily to accelerate the part of the machine that is already instrumented and already fast, while the unmeasured part quietly sets the total. The steering wheel is still broken. We have just fitted a bigger engine.
If this changed how you think about even one thing, the next post might too.




