Most teams have security scanning in the pipeline now. Dependency audits, a SAST pass, secret detection, an SBOM somewhere. The tooling problem is largely solved and the tools are mostly good.
The risk has moved. It is no longer that nobody is looking. It is that something is looking, reporting nothing, and nobody can tell the difference between a clean result and a broken one.
Six of twelve
We found this in our own work, which is the only reason we are confident about how common it is.
While authoring a set of code-quality and security rule packs, we ran twelve new rules for the first time. Six of them did not work. Not one threw an exception. Not one showed red. Every single failure rendered as a clean pass.
The reason is structural rather than careless. An empty result is ambiguous: "the analyzer looked and found nothing" and "the analyzer never ran" produce the same JSON. When a result is ambiguous, the default reading is the flattering one — and a broken governance tool grades every repository as excellent, so nobody goes looking.
A good score is the least investigated outcome in software. That is precisely what makes it the most dangerous place for a bug to hide.
The six shapes, so you can go and check yours
They are worth listing individually, because each one is a general shape rather than a one-off:
- Invalid YAML in a rule file. A pattern containing
:turned the entry into a mapping. The file did not load. The run succeeded. - A flag that means the opposite of what it reads like.
oxlint --silentsuppresses the findings, not the banner. We had been reading the banner. - A grammar with no pattern for a bare construct. You cannot match a lone
catch; the pattern needs the whole try/catch. Written the natural way, the rule can never fire. - The same problem for a bare attribute.
[IgnoreAntiforgeryToken]has to be attached to the declaration it decorates to be matchable. - An absence check the engine cannot express. semgrep cannot say "this file never mentions X". We had written a rule that asked it to. The fix is to delete the rule — never ship one that can never fire, because it will report a pass forever.
- A compiler-based linter with no working compile.
cargo clippyon a crate that does not build emits zero lints and exits happily.
All six were caught by the same discipline, and it is the single most valuable thing in this article: every rule needs a positive-control fixture containing the pattern it is supposed to find, run before the rule ships. A rule verified only against a real repository that happens to be clean has not been verified at all.
Read the tool's error channel, not just its findings
The second discipline is to stop treating a scanner's output as a list of problems and start treating it as a report on a run.
semgrep will tell you it could not load a rule file — in a notifications block, while still setting executionSuccessful: true and returning zero results. If you read only the results array, that is a clean scan.
The inverse is just as expensive. oxlint sets executionSuccessful from its exit code, so a run that found 53 genuine issues reports false. We treated that as a failed run and threw away every finding.
The rule that survives both: executionSuccessful: false means failure only when nothing came back. Per-file complaints never void a run. And a lens where every analyzer failed is not_run, never ok.
That last one has a frontend half, which we also got wrong. There was a single line in the interface rendering "Clean — no findings" for any empty lens. It undid the entire backend distinction in one string.
The same shape in a report, not a scanner
This generalises well past security, and the clearest example we have is from our Java upgrade path.
The verification step classifies build failures into three buckets: regressions we caused, pre-existing failures we inherited, and unproven. All three are keyed on the list of modules the build named as failing.
When Maven died before the reactor even started — an unreadable POM, an unresolvable coordinate — that list was empty. So all three buckets were empty. And the report read the emptiness as success: GREEN, the reactor builds, alongside a claim that the module had been repaired.
The build had not compiled a single line.
Any verdict derived from "did we find problems?" must first ask "did the thing run?" When several buckets share one input, that input being empty is a fourth state — not the intersection of the other three. Assert it explicitly, name it in the report, and exit non-zero.
The opposite failure: a confident wrong finding
Under-reporting is one half. The other half is a report that is loudly, authoritatively wrong, and it does more damage than it looks like it should.
A compliance report that declares HIPAA on a shopping cart does not get corrected. It gets the tab closed — and every true finding on that page dies with it. In this work precision is not a polish pass, it is the whole product.
Every precision bug we found in the compliance packs was found by running against a real repository. Not one was found by reading the pattern. Some of them:
mrnmatched inside a base64 integrity hash, and a Next.js storefront was declared a healthcare system. In the same family,rc4appeared insidesha512-pgRc4hJ4…and produced a broken-cipher finding.- Lockfiles are not code. Their hashes contain every three-letter sequence there is.
- Comments are not code. A line reading
# For FedRAMP compliance, set to 'sha256'conjured an entire regulated domain and a non-compliant-crypto finding. - A word in a log message is not personal data.
logger.exception("Failed to send export failure email")was counted as handling email addresses. - A URL is not an endpoint. An attribution link in a credits array was read as an outbound integration.
- Our corpus held only source files, so
SECURITY.md,.github/,.gitignoreand project files were invisible — and every "does this artifact exist" control answered from missing data, reporting present files as absent.
Two structural rules fell out of that work and we keep them:
A scoped "missing" check with no files in scope is not a failing control. "We could not look" and "it is absent" are different answers, and only one of them is a finding.
Signals are not equal. A flat count of matches missed a real payment integration that had exactly one strong signal, while inventing a healthcare domain from two weak ones. A payment SDK in your dependency list is a statement about what the system does. A word in a comment is not.
Dependencies: the part that ages while you sleep
Vulnerabilities get attention because they arrive as news. The quieter dependency risk is that your platform versions keep moving toward their end of support whether or not anyone is looking.
We treat this as its own family, scored separately. It reads what the repository declares — engines fields, .nvmrc, POM properties, target frameworks, requirements files, go.mod, Cargo.toml, composer and Gemfile entries, the Docker base image — and joins it against a curated, offline support matrix. It never invokes a package manager, so it works on an air-gapped checkout and cannot be changed by a registry.
Four details in that design are load-bearing:
- "Expiring" is computed from the end-of-support date at scan time, never authored. An authored status is wrong the day the date passes, and nobody notices for a year.
- "Restricted" is a licence change, not a vulnerability. When a widely used framework moved to a business source licence, no CVE feed modelled it, and it was a real problem for a lot of production systems.
- "Dead" carries a successor, because moving off it is a rewrite rather than an upgrade and should be planned as one.
- "Now" is the commit date of the tree being scanned, not wall-clock time. Two scans of the same commit agree with each other, which is what makes the score reviewable rather than a moving target.
Severity is not a work queue
A list of CVEs sorted by severity is not a plan, because severity describes the vulnerability and says almost nothing about you.
Scanners return more than an identifier and a rating. There is an exploitation probability, whether a fix exists at all, and the date that fix was published. Three different cuts fall out of that, and they call for three different responses:
- Actively exploited. High predicted exploitation, or in the top few percent. This is the queue, and it should escalate severity rather than sit inside it.
- Unfixable. No patch exists. Sorting it next to fixable items is a category error — the remedy is isolation, replacement or acceptance, not an upgrade.
- Exposure. A fix was published more than ninety days ago and has not been taken. This is the only one of the three that is a statement about your organisation rather than about the ecosystem, and it is usually the one worth showing a board.
Reachability changes the order
Two identical SQL injection findings in the same repository are not equally urgent. One sits in a controller on a public HTTP route. The other sits in a file nothing reaches from any entry point.
Because the structural pass already knows the routes and what is reachable, each finding can be tagged with its exposure — on the attack surface, internal, or in dead code — and ranked accordingly. It is the same class of bug and a completely different Tuesday.
This is also the cheapest honest improvement available to most teams: you probably already have both halves, and are just not joining them.
Show what was not checked
The last piece is a coverage view, and its value is entirely in the blanks.
Laid out against the OWASP Top 10, some categories are not reachable by any static scanner — insecure design is the obvious one. That should be stated, not left as an implied zero. Everything else falls into two groups: covered, and reachable-but-not-covered. The second group needs to say NOT CHECKED in as many words.
A blank you can see is a decision waiting to be made. A blank that renders as a green tick is a decision somebody has already made on your behalf, without telling you.
The question to ask of your own pipeline
Not "what did the scan find". That is the question the report is designed to answer.
Ask instead: which analyzers actually ran on the last build, which ones failed, and how would you know? Take one rule you rely on, write a file that should trip it, and push. If nothing happens, you have learned something more useful than any dashboard was going to tell you.
Security will not go wrong for you because a scanner was missing. It will go wrong because one was green.