2 · How a codebase is read

How the workbench reads your code

One pass over the repository produces everything the rest of the product uses. This section covers what that pass works out, how it is stored, how it decides which code can be reached, and the views you get on top of it.

The structural pass

One analysis runs over the repository. It writes a set of JSON files, called shards, that describe the code from different angles: what each file declares, which function calls which, where the program starts, what it asks the database for, which HTTP routes exist, and so on.

Every screen in the product reads those shards. The views below, the governance score, the Modernize screen and the planning step of any run all work from the same pass. None of them re-reads your source.

The Structural Intelligence view. The left rail lists the shards by group with their size on disk. The right pane shows one of them, either formatted or as raw JSON.
The Structural Intelligence view. The left rail lists the shards by group with their size on disk. The right pane shows one of them, either formatted or as raw JSON.

The same code gives the same result

The overview shard carries a merkle_root, a checksum over the tree that was analysed. The same source always produces the same value, and changing one file changes it. That is what lets the analysis be cached and compared. A later run can tell what actually moved instead of redoing everything, and two machines looking at the same commit agree on what they are looking at.

It sits in your working tree

The shards are written to .cognidev/ next to the code. You can read them formatted or as raw JSON, filter them by field or value, and commit them with the repository if you want your team to have them.

Watch this part of the tour

What the pass produces

The shards are grouped the way the product groups them. Which ones you get depends on the repository. A project with no SQL has no schema shard, and the polyglot and test shards only show up where they apply.

GroupShards
OverviewStack, language, counts, merkle root
Code StructureFiles, with what each one declares · Call Graph, with the calls it could resolve · Types, as UML · Structural Graph, with resolved calls, writes and handles plus infrastructure
Entry & ReachEntry Points · Reachability · Hot Files
DataEntities and the tables they own · Database call sites · SQL schema from the DDL
Behaviour & FlowsFlows, from an entry point to a side effect · State Map · Sequences · HTTP Routes · Roles and authorization · Lifecycle hooks
ArchitectureLayers · Domains · Folder summaries · Dependencies · Graph insights, meaning centrality, cycles and clusters · Modules, per runtime
Packages & FrameworkFramework API, with the real package signatures
Test AutomationSelenium suite: page objects, selectors, data injection
Product & DocsUse Cases · Reading Guide · Logs · Screens · Locator, a search index for finding where something is · Repos · Activity

Each row shows how big the shard is, which is a quick way to see how much of that thing the repository actually has. A large call graph next to a tiny hot-files shard describes a very different system from the other way round.

Entry points and reachability

Two shards do more than their size suggests.

Entry points are the places nothing else calls, which is where execution actually starts. Reachability walks the call graph out from those points and marks what can be reached.

That is what makes the dead-code count trustworthy instead of a guess. A file with nothing pointing at it is not necessarily dead, because it might be an entry point. A file that is referenced can still be unreachable, if everything referencing it is unreachable too. You need the entry points and the graph together to answer the question, and both are worked out from the code rather than guessed from how things are named.

The same walk produces hot files, ranked by how many calls come into them. Between them, the three tell you where to start reading, what is safe to delete, and what will be expensive to change.

Evidence and confidence

Findings come with the thing that produced them. The Tech Stack view shows this most clearly.

The Tech Stack view. The runtime is stated with how many files it was read from and a confidence figure. Each framework below opens to show the evidence.
The Tech Stack view. The runtime is stated with how many files it was read from and a confidence figure. Each framework below opens to show the evidence.

The runtime is given as .NET 10.0, from 20 source files, at 90 per cent confidence. The confidence figure is part of the finding, so you can decide how much weight to give it.

Under each framework is the evidence. content_contains:Microsoft.AspNetCore means that text was found in the source. file_present:…/Acme.Banking.csproj means that file is on disk. Neither is an opinion, and you can check both by hand in a few seconds.

The same page shows the key that decides which runs get offered for this repository, here aspnetcore. So you can see why a particular set of runs shows up later, at the point the stack was decided.

What it could not read

The second rule is about what the analysis could not do.

The Schemas view. It lists every signature and response shape that was captured, and any call site whose shape could not be worked out is listed too, marked as a gap.
The Schemas view. It lists every signature and response shape that was captured, and any call site whose shape could not be worked out is listed too, marked as a gap.

The Schemas view lists every signature and response shape the pass captured. The row worth looking at is at the bottom: a call site that is real, but whose shape the analysis could not work out. It is marked as a gap, and the screen says plainly that a gap is not the same as the call site returning nothing.

This matters more than it sounds. A tool that quietly drops what it failed to parse produces a clean-looking report that covers less of the system than you think. Here you can see what was missed, so a plan built on the analysis carries the same caveat, and you know which parts of it rest on complete information.

The same rule applies to runs

When a run cannot confirm one of its own checks from what it read, it says so instead of marking it passed. There is an example in Upgrading a .NET application in place.

Finding your way around

Understand it gives you a set of views rather than one report. Each card shows the count it found, so you can see the shape of the repository before opening anything.

The views. Each card names what it shows and the count it found. The counts come from the pass, so opening one is instant.
The views. Each card names what it shows and the count it found. The counts come from the pass, so opening one is instant.

These four answer the first question anyone has on a codebase they do not know: what is this, and where do I start reading?

ViewWhat it shows
Guided tourThe files in a sensible reading order, in chapters. Where to start, and what to read next.
DocumentationA full documentation site built from the analysis and written to .cognidev/docs/. No model is involved, so it costs nothing and cannot drift away from the code it was built from.
Tech StackRuntime, frameworks and versions, each with its evidence and a confidence figure.
Architecture tourThe architecture one layer at a time.

The two tours are different on purpose. One walks the files in reading order, the other walks the layers. On a repository you do not know, start with the first.

What the system does

These views describe what the system does, rather than how it is arranged.

ViewWhat it shows
DomainsThe business areas in the code and what sits in each.
Use CasesWhat the application does, in plain English, with the files that do it.
FlowPaths from an entry point through to a side effect.
SequenceCall sequences, as diagrams.
State MapModule-level variables and where each one is read or changed.
LifecycleHow each entity moves through its states. This only appears when at least one entity has a status to move through, so an application without one never shows it.

Use cases in more detail

This is the view that reads the system as a product rather than a program.

The Use Cases view. Each one is broken into the steps it performs, and each step names the files that implement it.
The Use Cases view. Each one is broken into the steps it performs, and each step names the files that implement it.

A retail banking core comes out as eleven use cases. Each is broken into the steps it actually performs. A transfer is submitted, screened for fraud, debited, credited, recorded, and the customer is told. Each step names the files that do the work.

That helps with two jobs. Someone new starts from what the system does rather than from the folder names. And when an audit asks where a business rule lives, you have a file to open instead of a search to run.

Contracts and data

These views cover what crosses the boundary and what the data behind it looks like.

ViewWhat it shows
RoutesEvery route: what goes in, what comes out, and what it touches.
Data ModelTyped fields, the values each one can hold, and the test data those values imply.
SchemasOpen a file and see every object, payload and schema declared inside it.
Stored ProceduresEvery routine: its signature, what it declares, the tables it touches, and what it does. Shown where the repository has stored routines.
Test DataEvery typed input, and the cases each one has.

Data Model and Schemas show the same information from opposite ends. Data Model starts with a type and tells you what it holds. Schemas starts with a file and tells you what is declared inside it, which is the question you have when you are looking at one. Which is more useful depends on what you already know.

Test Data exists because a typed input with known limits already tells you what its cases are. The view does not invent values. It works out what the declared types allow.

Structure and risk

These views cover how the code is wired together, and where it is likely to hurt.

ViewWhat it shows
UMLClass diagrams from the types that were parsed.
InsightsHotspots, cycles and clusters, from the call graph.
WiringRoutes, infrastructure, and how far a change can reach. Needs a structural graph for the runtime, which today means Go, Java, JavaScript, Python and Rust.
MapBrowse the files and see what flows in and out of each one.
DependenciesThe outside packages you use, and where each one is used.
Dead codeFiles nothing can reach, worked out from the entry points and the call graph.
TestsThe shape of the test suite itself. Shown where the project has a Selenium suite.

How far a change reaches

Insights and Wiring answer the question that decides how big a change really is. Insights ranks files by how central they are, so the file everything depends on gets named up front instead of found later. Wiring follows the call graph out from a change to show what it can reach.

Empty is an answer

A view with nothing in it says so in its own words: "Nothing unreachable — all files are reachable", or "No Selenium suite in this project". An empty result and an unsupported one never look the same.