What AI Context Limitations Teach About Software Development
Ask an AI coding assistant to fix a bug in one file, and there's a decent chance it patches the symptom at the call site rather than the logic several layers away that's actually causing it. That originating logic never made it into the context window for this particular prompt. Ask it to add error

Ask an AI coding assistant to fix a bug in one file, and there's a decent chance it patches the symptom at the call site rather than the logic several layers away that's actually causing it. That originating logic never made it into the context window for this particular prompt. Ask it to add error handling to a new module, and it may quietly invent a new convention rather than use the one the codebase already settled on two months ago. The file where that convention lives wasn't part of what it was shown this time. Ask it to add a helper function, and there's a real chance it writes a third, slightly different version of a date formatter that already exists twice elsewhere in the codebase. Not out of carelessness. It never saw the other two. None of this is primarily a training problem. It's a visibility problem. The model is only ever reasoning over what's in front of it. Even when a codebase is technically small enough to fit inside a large context window, there's now good evidence that long-context models don't attend to everything in that window with equal reliability. Information buried in the middle gets used less faithfully than information at the start or end. So even "just give it the whole file" doesn't reliably produce the thing a human means by "the assistant understands the codebase." What it produces is closer to a series of locally coherent, globally unaware decisions. Each one internally sound. None of them checked against the others. That failure has a name, and it's worth sitting with before reaching for the obvious fix. The honest reveal here isn't that AI has a context problem software development didn't already have. It's that software development has had exactly this problem for decades, and the industry never diagnosed it as a context problem. The actor doing the pattern-matching was human, and humans don't announce their context window the way a token limit does. Large COBOL systems became notorious for a reason that has nothing to do with the language. Nobody could hold the whole system in mind anymore. Nobody could say with confidence what depended on what, and eventually nobody dared touch the code. Object orientation was supposed to fix this. Java, and later C#, arrived with the promise that encapsulation and class design would finally make large systems reasoned-about rather than merely operated on. But there weren't enough developers who understood object orientation deeply enough to do that reasoning. A generation was trained quickly instead: bootcamps teaching the syntax of controllers, services, and dependency injection without ever teaching what a responsibility is or where it belongs. The result was code that used classes syntactically while remaining procedural in structure underneath. This is the anemic domain model, named as such over twenty years ago, and still the default shape of a great deal of enterprise Java and C# written today. That developer wasn't limited by a context window. The code was sitting right there, fully readable, no retrieval problem at all. What was missing wasn't access to information. It was the skill of holding a coherent model of the domain and checking new work against it. Handed a new requirement, that developer reaches for the nearest template: Controller here, Service there, Repository underneath. The same template gets applied every time, regardless of whether this particular responsibility actually belongs in that shape. It is functionally the same act an AI agent performs when it pattern-matches a new prompt onto a scaffold it has seen many times in training. Select the nearest known shape, apply it, move on. The causes are different. One agent literally cannot see the rest of the codebase. The other was simply never taught to look. Both arrive at the identical output: locally plausible code that was never checked against a standing model of the whole, because no standing model was ever being held in the first place. This is the distinction actually worth naming, because it's the one that survives contact with every objection. There is a real difference between applying a known pattern and designing. Pattern application is recognizing that a situation resembles something already seen and reaching for the matching shape. Design is discovering, through friction with something that won't simply agree with you, where a responsibility actually belongs. That friction might come from an ambiguous domain expert, or from code that gets uglier the further a wrong assumption is pushed. Design means revising the model when the friction shows the current one is wrong. A framework-trained developer applying Controller-Service-Repository to everything is doing the first. An AI agent completing a prompt by reaching for the nearest scaffold in its training is doing the first. Neither is doing the second. The reason is the same in both cases: neither is holding, or being made to hold, a single, standing, revisable model of the actual problem, checked against something outside itself that's allowed to disagree. This is where Fred Brooks is useful, not as decoration but as the vocabulary that already explains the failure. In No Silver Bullet, Brooks split software difficulty into two kinds. Essential complexity is the irreducible difficulty of the problem itself: the actual business rules, the actual domain logic, the things that would still be hard with a perfect language and infinite time. Accidental complexity is everything layered on top while trying to manage it: the frameworks, the deployment topology, the infrastructure. Essential complexity is inherent to the problem. Accidental complexity is a byproduct of the tools chosen to solve it. Design, in the sense just described, is the act of making essential complexity legible. It pins that complexity to a structure where it can be checked, disputed, and corrected. Pattern application, by contrast, doesn't engage with essential complexity at all. It substitutes a known shape for the work of discovering what the actual complexity is. That's exactly why it produces code that runs, passes its tests, and is quietly wrong relative to everything built around it. In a small application, none of this matters. The functional context is small enough that one person can hold the whole model in working memory at once. Even pure pattern application, or pure procedural code with no model at all, rarely goes wrong in a way anyone fails to notice. There's nowhere for a contradiction to hide when the whole system is visible at a glance. This is also why AI looks unreasonably good on small projects and demos: there's nothing to design that a known template doesn't already cover, so pattern application converges on something correct almost by default, and neither a human's context limit nor an AI's is ever put under real strain. The trouble starts once the system outgrows what a single mind can hold. At that point, essential complexity has to be made explicit. It has to be encoded into a structure with exactly one home per responsibility, or contradictions accumulate invisibly, because nobody is looking at all of it at once anymore, and nothing is forcing the parts to answer to each other. That's also, not incidentally, the actual argument against deciding the system's boundaries before anyone has written a line of code. Splitting an application into services upfront is a Big Design Up Front decision, made at the exact moment the least evidence exists about where the real boundaries are. Agile development's whole case against waterfall was that a model's correctness is discovered through contact with implementation, iteratively, and gets revised as that contact reveals it was wrong. Fixing service boundaries before that contact happens locks in a decomposition based on guesswork. It also makes the eventual correction far more expensive than it would have been inside a single codebase. "We drew this boundary in the wrong place" now means data migration and coordinated multi-team redeploys, instead of moving a method between two classes in one commit. A team can run two-week sprints, hold retros, call itself agile in every visible process, and still be doing waterfall at exactly the decision that matters most. The one place iterative revision was needed most was fixed before the iterating ever began. Once this is named, the shape of the rest of the argument follows. Humans have their own version of the AI's context boundary. It is cognitive rather than measured in tokens, and because nobody could point at it the way a missing file can be pointed at, it never got recognized as the same problem. Microservices are, in large part, what an organization reaches for once it senses the loss of overview without being able to name it. Shrinking what any one team has to hold looks like a fix for the mess. It isn't. It's a control mechanism against spaghetti, not a solution to it, and it works by hiding the symptom rather than restoring the visibility that was actually lost. That's the same move an AI agent's narrow context window makes by default: not removing the rest of the system, just not looking at it right now. The fix, in both cases, was never a bigger window. It's giving essential complexity an explicit, persistent, functional home that survives past whichever mind, or whichever prompt, happens to be looking at it at the time. That's what a rich domain model actually does, and it's the thread the rest of this argument follows. None of this is an argument that splitting an application is never justified. It's an argument for being precise about what problem the split is actually solving, because the honest justification is rarely the one people reach for. The most common defense of microservices and event-driven architecture isn't really "it reduces complexity." Most engineers who've worked with both know better than that by now. The more common defense is that it handles load and concurrency: services can scale independently, message queues smooth out spikes, producers don't block on slow consumers. That's a real category of benefit, and it deserves to be taken seriously rather than waved away. But look closely at what actually delivers each of those benefits. Temporal decoupling, a producer not blocking on a slow consumer, comes from asynchronous messaging, not from splitting into separately deployed services. A queue sitting inside a single application, backed by a single database, delivers the same decoupling without a network boundary, without a saga, without an eventual-consistency problem. There's still one deployable and one source of truth. High concurrent load, thousands of simultaneous in-flight requests, is handled by non-blocking I/O and reactive programming within a single process, and by horizontal scaling of that same stateless process across more instances. None of that requires decomposing the domain into separate services either. For request-serving application code specifically, once the process is stateless and the database isn't the bottleneck, horizontal scaling doesn't require splitting the domain into separately deployed services at all. If the database is the bottleneck, splitting the application layer into services does nothing about that. The fix is partitioning the data, which is an orthogonal decision entirely. So when a system has genuinely been split for load or concurrency reasons, it's worth asking a follow-up question. Which specific mechanism is doing the work: async I/O, horizontal scaling, a queue? Did any of those actually require splitting the domain into separate services with separate databases, or would they have worked identically inside one deployable? Most of the time, the mechanisms that solve load and concurrency don't need the split at all. What they need is asynchronous handling and horizontal scaling of the application layer, both of which are available inside a single, well-modeled system. There is a narrower, real exception, and it's worth being precise about it rather than dismissing it along with the rest. Companies like Netflix, Uber, and Amazon had components with such different resource profiles, video streaming against billing, ride-matching against fraud detection, that scaling them together meant chronically over-provisioning one to cover the other's peaks. Splitting there is solving heterogeneous resource-scaling economics. It isn't complex, interrelated domain logic. It's comparatively simple logic executed at extreme, uniform volume, with real regional variance in how it needs to run. That's closer to a strategy pattern deployed at application scale than a fix for entangled business rules. It is nearly the inverse of the problem most enterprises actually have, which is moderate volume with deeply interrelated logic that benefits far more from staying coherent in one model than from being split and independently scaled. The pattern most cargo-culted across the industry was built to solve a problem shaped almost the opposite of the one most of its imitators actually have. Which leaves the real question for any given split: was it made because a measured, Netflix-shaped constraint on load or resource economics actually existed? Or because splitting into services is what's perceived as the industry standard, the thing a competent architecture is simply assumed to look like? The saga is a good place to see this cost made concrete, because it is a clean example of the trade the split actually requires. It is a technological mechanism, coined in a 1987 database paper, rather than something a domain owner is likely to request by name. No one asks for "a saga"; they ask for "cancel the order." That doesn't mean sagas are never warranted. A warehouse order where a picker is already walking a shelf to retrieve a product, and that product then needs to be reassigned mid-pick, is a genuine physical process crossing a genuine system boundary. Something saga-shaped, or at least a compensating action, is doing real work there. But the fact that it's sometimes warranted doesn't make it cheap. It's encoding, by hand, in application code, a rollback guarantee that a single JDBC transaction gives away for free whenever the boundary being crossed is only a database, not the physical world. Every saga is a bet that the coordination problem is real enough to justify trading away that guarantee. The more a system can stay inside one transaction, the fewer failure scenarios anyone has to understand, predict, or defend against later. The failure mode this produces is also strictly harder to catch than the monolith's. A coupling bug in a single well-modeled application surfaces at compile time or in a test run. The same coupling, once hidden behind a network boundary and an event bus, surfaces as misaligned data discovered by an analyst two years later, with no stack trace and no one left who remembers why the two systems were ever related. The COBOL-era fear, that no one understands what depends on what and no one dares touch the code, isn't fixed by this architecture. It's given a longer fuse. Microservices hide contradiction behind a network boundary. AI-assisted development hides it behind generation speed. The mechanism is different. The failure is the same one described at the start: essential complexity stops being checkable against a single, coherent model, and the cost doesn't disappear. It moves downstream, to whoever finds the mismatch later. An AI agent, as currently built, is a "you say, I do" tool. It generates code in response to an instruction and checks the result against a criterion, a test, a spec, iterating until that criterion is satisfied. That's a real and useful loop, but notice precisely what it verifies: whether the implementation conforms to the instruction it was given, not whether the instruction is a coherent extension of everything decided before it. That's verification, not design. Design requires discovering whether the criterion itself is still the right expression of the problem, and a domain expert supplies something an agent's own generated tests cannot: contradiction from outside the model. Saying something today that doesn't match what was said last month. Revealing, through the process of implementation, that two things once described as separate are actually the same concept. That contradiction is what forces a model to be revised rather than merely satisfied. Absent a standing model to check against, an agent has no equivalent source of friction. It will produce a second, contradicting version of already-settled logic exactly as confidently as it produced the first. It's the identical anemic, transaction-script failure the industry has produced every time responsibility wasn't pinned to one checkable location, just now generated at a volume and speed no bootcamp graduate ever came close to matching by hand. A correctly built domain model earns its keep twice over here. It's not only what makes a large system maintainable. It's what would let an agent's output actually be checked for contradiction, mechanically, the same way it lets a human reviewer catch it. If a given responsibility has exactly one home in the code, a conflicting instruction doesn't quietly coexist somewhere else. It collides, visibly, with what's already there. Without that structure, there is no fixed point for a contradiction to collide with. Nothing, human or AI, has anything to check new output against except the same loose criterion that let the first contradiction through. Put together, the pattern across four decades and every "next big architecture" is the same one. In a small enough system, one mind can hold the whole model, so pattern application, procedural code, and AI-generated code all tend to work. There's nowhere for a contradiction to hide. Past that point, essential complexity has to be made explicit. It has to be encoded into a structure with exactly one home per responsibility, or contradictions accumulate invisibly until the cost of finding them exceeds the cost of ever having built the system. Every subsequent attempt to manage that growth without doing the actual work of design, pre-split service boundaries, event buses reached for as a default, unscaffolded AI agents pattern-matching against a context window that doesn't include the rest of the system, doesn't shrink the field. It hides the seams, trading a fast, visible failure for a slow, invisible one. The monolith, correctly modeled, isn't a legacy pattern waiting to be modernized away. It's the fastest available warning system for exactly the failure everyone is trying to avoid. Messy, inconsistent, contradictory logic gets spotted immediately, because there's only one place for it to be, and everyone touching that place sees what's already there. None of this is an argument against using AI. It's an argument against using it as a shortcut around design. Pointed at a ticket and asked to produce code, AI is very good at exactly the failure mode this article has been describing: pattern application, at speed, with no standing model to check itself against. Used as a discussion partner instead, something closer to a fast, tireless consultant to think a model through with, it's a different tool doing different work. Its value there isn't in writing code. It's in supplying another pass of friction: surfacing known shapes, known failure modes, known alternatives, fast enough to test a model against them before a line gets written. And a correctly built domain model doesn't actually need much code once that friction has done its work. Much of what fills a typical enterprise codebase was never essential complexity to begin with. It's the scaffolding built to manage its absence: dependency injection wiring, service interfaces with exactly one implementation, layered DTOs and mappers whose only job is to cross a boundary that didn't need to exist, saga orchestration standing in for a transaction that was given away for free. Strip that away, and what's left, the actual domain objects and the connections between them, is small in comparison. Generating that was never the hard part, for a human or for AI. Finding it was, and no amount of generation speed substitutes for that. Which points at the actual subject of this article, and it was never AI's context window. That was only the instrument that made the problem visible, because it fails in a way that can be pointed at: these files were shown, those weren't. The real limitation was always the human one, and it doesn't go away once AI is in the room. Below a certain size, none of this matters, because pattern application converges on something correct without anyone needing to hold more than fits comfortably in view. Above it, the limitation is permanent, and it isn't solved by a longer conversation. Outsourcing the design work to AI is closer to reading a correct, well-written course on pottery and mistaking that for having thrown. The book can be right that uneven pressure collapses a thin wall. It cannot give the moment where a particular wall actually wobbles under a particular pair of hands, and the rule stops being words and becomes something felt. A domain expert, or a codebase's own resistance, produces exactly that kind of friction: temporal and lived, a decision made, watched for months, discovered wrong only once its consequences actually showed up. AI, used as a discussion partner, can sharpen a model that's already being formed this way, surfacing a known failure mode before it gets rediscovered the hard way, the way a good book can tell someone in advance that thin walls collapse. It cannot supply the throwing itself. That still has to be lived, by someone, over time, inside the actual domain, and nothing currently called AI does that. Brooks was right that there is no silver bullet. The bullet was never going to make essential complexity disappear. The only thing that has ever actually worked is not losing sight of it, and that has always been a human limitation before it was anyone's context window. Forty years of trying anyway have just gotten faster at showing the bill.
Key Takeaways
- β’Ask an AI coding assistant to fix a bug in one file, and there's a decent chance it patches the symptom at the call site rather than the logic several layers away that's actually causing it
- β’This story was reported by Dev.to, covering developments in the dev space.
- β’AI advancements continue to reshape industries β read the full article on Dev.to for complete coverage.
π Continue reading the full article:
Read Full Article on Dev.to βShare this article



