There are two different jobs hiding inside every modernization project.
The first is to move the system: same behaviour, new runtime, new framework, new place. The second is to make it better: fix the shape, adopt the target framework's idioms, drop the things that were always wrong.
Both are legitimate. Doing them at the same time is where migrations go badly, and it goes badly in a specific way that is worth naming.
Improvement is the thing you cannot verify
When a migration only translates, you have a test available to you: the old system's behaviour. Every question of the form "is this right?" has an answer somewhere in the source, and disagreements are settled by looking.
The moment the migration also improves, that test is gone. The new code is deliberately different from the old code, so a difference is no longer evidence of a bug. Every review becomes a judgement call, every judgement call takes a senior engineer, and the project's throughput collapses to the rate at which those people can read.
This is not an argument against improving the system. It is an argument for sequencing: translate to something you can check, then improve against something you can run.
What this cost us
We learned the sharp version of this from our own tooling, on the Selenium to Playwright translator.
Playwright's documentation puts user-facing locators first — getByRole, getByLabel — and that advice is correct. So our translator turned By.id("email") into getByLabel('Email'), which is exactly the sort of upgrade a modernization is supposed to deliver.
It was fabrication. Playwright's advice is written for a person looking at a rendered page, where the accessible name comes from the DOM — a <label>, an aria-label, the element's own text. A Selenium source file contains locator strings and nothing else. There is no accessible name in it to read.
The delivered suite matched zero elements and could not run. Roughly 187 locators were invented, twelve files were accidentally correct, and the report described the result as "hardened".
Any transform that upgrades a construct toward the target's best practice must be able to point at the evidence in the source. If the evidence lives somewhere the migration never sees, the upgrade is a guess wearing the costume of a decision.
The rule we now hold ourselves to: translate to provable equivalents. By.id(x) becomes #x. By.className(x) becomes .x. CSS and XPath go across verbatim, with the original kept as a comment so the audit trail survives.
There is exactly one evidenced upgrade in that translation, and it is worth understanding why it qualifies: By.linkText(t) becomes getByRole('link', { name: t }), because a link's visible text is its accessible name. The evidence is inside the locator itself. Everything else waits for a step that can see the running application.
We also deleted a questionnaire option. There had been a setting for locator strategy, defaulting to "resilient first". The question was the bug: no answer to it was ever going to be right, so offering it only distributed the blame.
The green build proves less than you think
The other half of the problem arrives at the end. A migration finishes, the target compiles, the pipeline is green, and somebody has to decide whether it is done.
Most of what a modernization can check at that point is target-only. Does it build. Do the services agree with each other's contracts. Do the generated tests pass. All useful, and none of it says anything about the source. Nothing in a green build asserts that a behaviour in the new system reproduces its counterpart in the old one.
That gap is why we ended up building comparison as a first-class thing rather than a report at the end.
Comparing two repositories
The model is simple enough to state in a sentence: equivalence is a structural diff of one parser's output over two inputs.
The workbench already runs a structural pass over a repository — the call graph, the routes, the entities, the tables, the modules, the dependencies. To compare two systems, we run that same pass on both sides and diff the results. Not two analyzers agreeing. One analyzer, twice.
That matters more than it sounds. Two different tools reading two different codebases will disagree for reasons that have nothing to do with the codebases. One tool run twice can only disagree about the inputs.
In the product this appears as a Modernize option, "Compare against another repository". You pick the second repository from your connected accounts, it is cloned into a workspace library outside the project being analyzed, and both sides get the same treatment. The output is a report written into the source repository, so it can be committed and reviewed like anything else.
What the comparison decides
Overlap is measured across five vocabularies — routes, entities, tables, domains and dependencies — and that overlap decides the shape of the work:
- Migrate, at roughly 40% overlap or more. These are two versions of the same system. One should become the other.
- Selective, between about 15% and 40%. They share real ground but neither subsumes the other. Move specific capabilities, not the repository.
- Merge, below about 15%. Whatever the two do, they do not do the same thing, and treating one as the target for the other will hurt.
A separate layer-by-layer scoreboard decides direction — which of the two should be the base. Size is measured and deliberately never scored. The larger repository is not the better one, and a metric that quietly assumes it will always recommend keeping whichever system has accumulated the most code.
The rules that keep it honest
Two of them do most of the work, and both are about refusing to score an absence.
A vocabulary that one side is empty in does not get a vote. A scheduling library that serves no HTTP routes is not in 0% agreement about routes with a web application. It has no opinion about routes. Averaging that zero into a similarity score produces a confident number about nothing.
A layer with missing evidence on either side is "not decidable", never ranked against absence. "We could not see it" and "it is not there" are different answers, and only one of them belongs in a recommendation.
Why we verified it on a port
The first pairs we tested with were two independently written systems in the same business — an e-commerce application in C# against one in Java. That kind of pair can tell you the tool produced a number. It cannot tell you the number is right, because nobody knows what the right answer is.
So we ran it against a genuine port: Quartz.NET against Quartz, the same scheduler written twice in two languages. The same class names, the same methods, the same QRTZ_* tables on both sides. Here the correct answer is knowable independently, so a wrong reading is visible. It came out as migrate, at 58% overlap with 84% containment, which is the right shape for a port.
Getting there surfaced six scoring defects and three gaps in what we were capturing that four earlier pairs had not. Two are worth passing on, because they will break anyone's cross-runtime comparison:
- The .NET interface prefix.
ISchedulerandSchedulerare the same concept, and treating them as different names hid 20 types and 203 callables — which happened to be the public surface of the library. The entire comparison was reading the wrong half. - Letter case in SQL.
VARCHARagainstvarcharturned three real column differences into twenty-three.
Both are folded now, and both spellings are printed in the output, because a normalisation you cannot see is a normalisation you cannot check.
Every time we changed the scoring, one repository in the set was kept as a control whose numbers were required to stay exactly the same. Without one, a fix that improves the pair you are looking at and quietly breaks four others looks like progress.
Where this leaves a modernization
The practical version of all this is short.
Decide up front which of the two jobs you are doing, and do not let the second one leak into the first. Translate to equivalents you can point at. Keep the original visible in the output, so somebody can check the translation without going back to the old system. Then, with something that runs, improve it against reality rather than against documentation.
And when you want to know whether the move actually landed, compare the two repositories with one analyzer and read the layers it could not decide, not just the ones it could.
The comfortable failure in this work is not a migration that breaks loudly. It is a migration that finishes, compiles, and quietly does something slightly different from the system it replaced — with a report that says hardened.