← All cases

03 · Case study

Governance review of a vibe-coded application

An application a model wroteAny language · report first, then fix

The system

Detector families
12
Questions asked
38
Default mode
read-only
Model reviews
the residual

The problem

The application works. It was written quickly, file by file, by a model, and it does what it was asked to do. What it does not have is a shape. There are components no one will open twice, comments that restate the line below them, four copies of the same fetch helper, folders that follow no layout, stub functions that return the happy path, a key committed in a config file, and an import of a package that appears in no manifest. Each of those is small on its own. Together they are the reason nobody on the team can say what changing one file would cost — and the reason the next model that touches it makes the same code worse.

What ran

  1. Build the structural map first

    Files with per-symbol line spans, entry points, reachability, HTTP routes, hot files. The scan reads that map, so a god-file is counted rather than estimated, and the same finding ranks differently depending on whether the file is reachable from a route or reachable from nothing.

  2. Scan, deterministically

    Twelve families: god-files and mega-functions, over-commenting, template CSS, code written against no house style, folder-structure drift, copy-paste, placeholder stubs, chat artifacts, fake data, hardcoded secrets, undeclared dependencies, risky code. No model call. The same repository produces the same findings.

  3. Rank by exposure, not by count

    A duplicated block behind an authenticated route is not the same finding as a duplicated block in a script nothing imports. The structural map is what tells them apart, so the top of the list is worth reading.

  4. Review the residual

    Optional, and fed only the small ranked set: confirm or dismiss each finding, and surface what a scanner cannot see — an API used in a way that compiles but is wrong, a function that does nothing, a structure built for a requirement that never arrived.

  5. Fix it, on a branch

    Remediate mode applies the provably safe deletions directly, then marks every remaining finding in place and fixes it one file at a time — scoped to the marker, grounded in the structural evidence, compiled, and committed. Never a whole-file rewrite.

  6. Re-scan and compare

    The structural analysis and the detector run again over the fixed code. The before-and-after finding count is the proof, and it is computed rather than claimed.

What it did

What lands in the repository

ai-governance/AI-GOVERNANCE.md for the diff · a self-contained HTML report for the review · findings.json for whatever you build on top. Committable files in the repository, not artifacts hidden in a dot-folder.

Three fix scopes, chosen up front

Safe deletions only — banner comments and chat leftovers, no model, cannot change behaviour. Recommended — also rename generic identifiers, implement the stubs, extract the copy-paste, move secrets to configuration, replace fake data. Aggressive — also split the god-files, relocate the drifted files, normalise the template CSS. The highest churn, and the one to review line by line.

A gate you can put in CI

Pick the categories that block: any secret, any dependency that resolves to nothing in the registry, god-files over a threshold. The exit code is the signal. Placeholders go back to the fix loop, because that is what it is for; a secret does not, because a person decides where it goes.

Where it ended

A ranked, evidenced list of what is wrong, and — if you asked for it — a work branch where the safe things are already fixed, one commit per file, with a second scan showing the count that dropped and the count that did not. The report says in plain words what it is: engineering hygiene and evidence. It does not certify EU AI Act, NIST or ISO 42001 compliance, and it says so every time it runs, because a scanner that implies otherwise is the most expensive finding in the report.

What happens next

  • Run it read-only first. The list is the argument for doing anything at all, and it costs nothing to produce.
  • Remediate at the recommended scope, review the branch, then raise the scope if the diffs read well.
  • Wire the gate into CI at the categories you actually care about, so the next generated file is caught on the pull request rather than in a quarterly audit.

The reason this works is the ordering. Reading the code structurally before scanning is what turns a list of two thousand style complaints into a ranked list of the forty that sit on a route someone can reach. The model is never asked to find the slop — it is asked to judge the residual and to fix one marked seam at a time, which is a small enough job to compile-check.