03 · Case study
An application a model wroteAny language · report first, then fix
The system
The problem
The application works. It was written quickly, file by file, by a model, and it does what it was asked to do. What it does not have is a shape. There are components no one will open twice, comments that restate the line below them, four copies of the same fetch helper, folders that follow no layout, stub functions that return the happy path, a key committed in a config file, and an import of a package that appears in no manifest. Each of those is small on its own. Together they are the reason nobody on the team can say what changing one file would cost — and the reason the next model that touches it makes the same code worse.
What ran
Files with per-symbol line spans, entry points, reachability, HTTP routes, hot files. The scan reads that map, so a god-file is counted rather than estimated, and the same finding ranks differently depending on whether the file is reachable from a route or reachable from nothing.
Twelve families: god-files and mega-functions, over-commenting, template CSS, code written against no house style, folder-structure drift, copy-paste, placeholder stubs, chat artifacts, fake data, hardcoded secrets, undeclared dependencies, risky code. No model call. The same repository produces the same findings.
A duplicated block behind an authenticated route is not the same finding as a duplicated block in a script nothing imports. The structural map is what tells them apart, so the top of the list is worth reading.
Optional, and fed only the small ranked set: confirm or dismiss each finding, and surface what a scanner cannot see — an API used in a way that compiles but is wrong, a function that does nothing, a structure built for a requirement that never arrived.
Remediate mode applies the provably safe deletions directly, then marks every remaining finding in place and fixes it one file at a time — scoped to the marker, grounded in the structural evidence, compiled, and committed. Never a whole-file rewrite.
The structural analysis and the detector run again over the fixed code. The before-and-after finding count is the proof, and it is computed rather than claimed.
What it did
What lands in the repository
ai-governance/AI-GOVERNANCE.md for the diff · a self-contained HTML report for the review · findings.json for whatever you build on top. Committable files in the repository, not artifacts hidden in a dot-folder.
Three fix scopes, chosen up front
Safe deletions only — banner comments and chat leftovers, no model, cannot change behaviour. Recommended — also rename generic identifiers, implement the stubs, extract the copy-paste, move secrets to configuration, replace fake data. Aggressive — also split the god-files, relocate the drifted files, normalise the template CSS. The highest churn, and the one to review line by line.
A gate you can put in CI
Pick the categories that block: any secret, any dependency that resolves to nothing in the registry, god-files over a threshold. The exit code is the signal. Placeholders go back to the fix loop, because that is what it is for; a secret does not, because a person decides where it goes.
Where it ended
A ranked, evidenced list of what is wrong, and — if you asked for it — a work branch where the safe things are already fixed, one commit per file, with a second scan showing the count that dropped and the count that did not. The report says in plain words what it is: engineering hygiene and evidence. It does not certify EU AI Act, NIST or ISO 42001 compliance, and it says so every time it runs, because a scanner that implies otherwise is the most expensive finding in the report.
What happens next
The reason this works is the ordering. Reading the code structurally before scanning is what turns a list of two thousand style complaints into a ranked list of the forty that sit on a route someone can reach. The model is never asked to find the slop — it is asked to judge the residual and to fix one marked seam at a time, which is a small enough job to compile-check.