← Back to Blog

The Hallucination Bug We Filed Against Our Own Pipeline — And How We Fixed It

Mid-migration, our own tool generated a service that called a method which does not exist. The code compiled in the model's head and nowhere else. Here's the issue we opened, the root cause we traced, and the fix — which had nothing to do with a cleverer prompt.

We Caught Ourselves Hallucinating

This is not a post about how other people's AI tools hallucinate. It's a post about the day one of ours did — in a real migration run, on a real codebase, in a step we thought we had locked down.

The task was a .NET monolith being pulled apart into services. The pipeline had already done the boring, deterministic part correctly: it had parsed the source, built the call graph, and classified components by layer and domain. Then it reached the step where a model actually writes a new service class. The generated code was clean. It was idiomatic. It named a repository method, called it with sensible arguments, and moved on.

The method did not exist. Not in the target framework, not in the project, not anywhere. The model had produced a signature that looked exactly like something that should be there — and the build was the first thing to disagree.

A hallucination isn't the model being stupid. It's the model being confidently average. It filled a gap with the most statistically plausible method name, and plausible was wrong.

We could have shrugged. The build caught it; a human fixed it in thirty seconds. But "the compiler will catch it" is not a strategy — it's a tax you pay on every single generated file, forever, and it only works for the errors a compiler can see. So we did the thing you do when a system misbehaves in a way that will recur: we filed an issue against ourselves.

Issue: model invents a repository method during service generation

Severity: high — silent architectural risk, not just a build break.
Observed: generated OrderService calls _repo.FindByCustomerAndStatus(...). No such method exists on the repository interface. Nearest real method is Query(Expression<Func<Order,bool>>).
Expected: generation should only reference methods, types, and signatures that provably exist in the target — or explicitly flag a gap for a human, never invent one.
Root-cause hypothesis: the generation step was given the task but not an authoritative inventory of what the target actually exposes. It guessed.

Why "Just Write a Better Prompt" Was the Wrong Answer

The first instinct — everyone's first instinct — is prompt engineering. Add a line: "only use methods that exist." Add a few examples. Turn the temperature down. We've watched teams spend weeks here.

It doesn't work, and it can't, for a structural reason. A prompt instruction like "only use methods that exist" is asking the model to check its output against knowledge it does not have. The model has never seen your repository interface. It has seen ten million repository interfaces from its training data, and it is answering with the average of those. Telling it to be accurate about a specific system it can't see just makes it more confident about the same guess.

The hallucination wasn't a prompt defect. It was a context defect. We had asked the model to generate against an API surface we had never actually put in front of it. The fix had to happen before the model was invoked, not inside the instruction we gave it.

0
Prompt tweaks that reliably eliminated invented signatures in our before-testing
100%
Of the hallucinated methods were "plausible" — sensible name, sensible args, wrong reality
1
Architectural rule that actually closed it: never let the model guess a fact we can compute

The Rule We Already Had — And Had Quietly Violated

We have a standard that predates this bug, and breaking it is exactly how the bug got in. We call it heuristic before LLM: if a fact can be derived from parsing the code or diffing the tree, it must not be left to a model. The model is the finishing layer, not the engine.

The list of methods a repository exposes is a fact. It is sitting right there in the source. There is no reason on earth to let a probabilistic system guess at it when a parser can state it. We had honored this rule for analysis — the call graph, the layer classification, the dependency map were all deterministic — and then quietly abandoned it at the one step where it mattered most: the moment of generation.

Determinism over magic. Two runs over the same code produce the same facts. The model is allowed to be creative about how it writes a service — never about what the surrounding system actually contains.

The Fix: Hand the Model the Facts It Was Guessing At

The fix is not a secret and it is not a trick. It's an ordering discipline. Before the generation step runs, we assemble a deterministic, parsed record of the units it will touch — what we internally call a Card — and put that in front of the model as ground truth. A Card is not prose. It's structured facts pulled straight out of the parser:

  • Signatures that exist — the real methods, params, and return types on every interface and type in scope. If a method isn't on this list, it isn't real, and the model has no cover for inventing it.
  • Calls and callers — what this unit invokes and what invokes it, resolved to actual identifiers, so generated code wires into the real graph instead of an imagined one.
  • Reads and writes — the data items and fields a unit actually touches, so behavior-preserving translation has something to preserve against.
  • Side effects — the database writes, I/O, and external calls that are really there, flagged deterministically rather than inferred.

For framework APIs — the .NET base class library, the NuGet packages actually referenced by the project — we go one step further and build a signature inventory from the real package metadata, not from the model's memory of what those packages "usually" look like. When the pipeline needs to know whether a method exists on a framework type, it reads the answer. It doesn't ask the model to recall it.

The ordering that closes the gap

Parse first. Extract the real signatures, calls, and data flow — no model involved. Ground second. Put those facts in front of the model as the authoritative surface it must generate against. Generate last. The model composes a service; it does not get to invent the world the service lives in. When something genuinely isn't in the Card, the correct output is a flagged gap for a human — not a confident guess.

Once the real signatures were in the context, the invented method simply stopped appearing. Not because we told the model to stop — because we removed the vacuum it was filling. The model was never trying to lie. It was answering a question we had failed to give it the data for.

What We Deliberately Did Not Do

We want to be precise here, because "how we fixed it" is easy to over-claim. A few things we specifically avoided:

We did not build a second, parallel path. The temptation with a bug like this is to bolt on a special-case checker just for method calls. We didn't. The fix flows through the same generation pipeline every other feature uses — one path, grounded better. Two modes that diverge become two bugs.

We did not lean on the model to self-verify. Asking a model "are you sure this method exists?" is asking the same unreliable oracle twice. Verification of a structural fact belongs to the parser and the build, not to a second opinion from the thing that got it wrong.

We did not turn this into prompt lore. There is no secret incantation in a system prompt doing the heavy lifting, and if there were, it would be the fragile kind of fix that breaks the next time a model version changes. The fix is architectural: the facts are in the context because we put them there deterministically, and that holds regardless of which model runs.

If your defense against hallucination is a sentence in a prompt, you don't have a defense. You have a wish. The durable fix moves the fact out of the model's memory and into the model's input.

The General Shape of Every Hallucination

Once we'd traced this one, the pattern generalized. Nearly every code hallucination we've seen — in our pipeline or anyone's — has the same skeleton: the model was asked to be authoritative about a specific fact it had no access to, so it substituted the statistical average and delivered it with total confidence.

An invented API method. A config key that doesn't exist. A call to a service that was decommissioned two years ago. A field name that's almost right. Every one is the same failure: a gap in grounding, filled with plausibility. And every one has the same fix: figure out which facts are deterministically knowable, compute them, and hand them over before the model writes anything.

This is why we keep saying the model isn't the bottleneck. The frontier models are extraordinary at composition when they're standing on solid ground. They're dangerous when they're improvising the ground itself. The engineering work — the part that's genuinely hard and genuinely ours — is making sure they never have to improvise a fact we could have just told them.

What Changed After the Issue Closed

Three things came out of this beyond the immediate fix.

First, we made the rule a gate, not a guideline. Generation steps now assemble their grounding Card as a precondition; a step that would generate against an un-parsed surface is treated as a defect, the same way a missing test is. The bug got in because the rule was aspirational at one step. Now it's enforced.

Second, we made the gaps visible. When something genuinely isn't derivable — a truly ambiguous case where the source doesn't settle it — the pipeline surfaces it as a flagged decision for a human instead of quietly guessing. A visible "I don't know" is worth infinitely more than a confident wrong answer, because you can act on the first one.

Third, it sharpened how we think about the whole product. Every place a model touches a real codebase, we now ask the same question first: what facts is it about to guess at, and which of those can we compute instead? The answer is almost always "more than you'd think." That question, asked relentlessly, is most of what separates a demo from a system you'd let near production code.

Parse First. Ground Second. Generate Last.

The honest version of this story is that we hit the exact failure we warn customers about, in our own tool, on a normal Tuesday. What we're proud of isn't that it never happened — it's that the fix wasn't a patch. It was the pipeline doing what it already claims to do, at the one step where we'd let it slide.

Hallucination is not a mystery and it is not unbeatable. It's a grounding problem wearing a scary name. Compute the facts you can compute. Put them in front of the model. Let it compose, not invent. And when it truly doesn't know, make it say so out loud.

That's the discipline behind CogniDev — not a better prompt, but a better order of operations. Parse first. Ground second. Generate last.

See What Grounded Generation Looks Like on Your Codebase

Get a free structural assessment of your system — the parsed call graph, dependency map, and API surface we'd ground generation against — before a single line of code is written.

Request a Free Assessment