Making a large codebase legible to AI
Generated code was plausible and wrong in the same ways every time. The fix was not a better prompt - it was giving the model the same context a new developer gets.
The first six months of using AI seriously on a large Shopify codebase went like this: the code it produced was syntactically fine, looked like it belonged, and was wrong in the same few ways every time.
It invented helper functions that did not exist. It reimplemented a snippet already in the theme. It named things in a style the team abandoned two years ago. It solved the problem in a way that worked in isolation and broke a convention that existed for a reason nobody had written down.
The instinct is to blame the model and write longer prompts. The longer prompts help a little and then stop helping.
The actual problem was that a new developer joining this codebase gets a week of context, a code review, and someone to ask. The model got a file and a sentence.
Context is a build artifact
What changed things was treating context as something you construct deliberately, in layers, rather than something you paste.
What exists. An accurate, maintained inventory of the sections, snippets and helpers already in the theme, with what each one is for. Most bad generations were reinventions - the model could not know a thing existed, so it built a worse version. This layer alone removed a large share of the noise.
How we do things here. Naming, file placement, the shape of a section, what goes in a snippet versus a section, which patterns are deprecated. The tacit knowledge a reviewer applies without thinking. Writing it down for the model turned out to be worth doing for the humans too, which was not the point but was a nice result.
Why it is like this. The constraints. Platform limits worked around, decisions taken for a reason that is not visible in the code. Without this the model confidently “fixes” a workaround that exists deliberately.
What this task touches. Task-specific: the relevant files, right now.
The first three layers are stable. They are maintained like documentation, because that is what they are. The fourth is assembled per task.
Then stop trusting it
Better context makes generated code better. It does not make it correct, and building a process that assumes it does is how you end up shipping confident nonsense.
Two things sit downstream of every generation:
Automated tests as the first reader. Not written for AI specifically - but the moment a large fraction of new code arrives from a generator, the test suite stops being a safety net and becomes the primary reviewer. Weak coverage plus fast generation is a genuinely bad combination, and it degrades quietly.
AI review against our own conventions. A second pass whose only job is to check the code against the written rules - naming, structure, the logic conventions the team agreed on. It is good at exactly this: mechanical, tireless, no ego. It catches the drift that human reviewers stop noticing on the fortieth pull request of the week.
Neither of these replaces a person reviewing intent. They remove the mechanical objections so that the human review is about whether the thing is a good idea.
What actually improved
Less time spent rejecting plausible-looking code. Fewer duplicate implementations of things that already existed. More consistency across work done by different people, because everyone - human and otherwise - was reading from the same written conventions.
And a genuine side effect: the codebase became easier for new developers, because everything that had lived in senior people’s heads was now written down. We wrote it for a model and the team benefited more.
The uncomfortable bit
This only works if the layers stay accurate. Context that has drifted from the code is worse than no context, because it produces wrong answers with high confidence, and confident wrong is exactly the failure mode you were trying to fix.
So it is maintenance. Somebody owns it. If nobody owns it, it decays, and in about three months you are back to writing longer prompts.