06 · Case study
A system of several servicesREST routes · SQL · stored procedures
The system
The problem
Most teams have two options for test data and both are bad. Someone types twenty rows once, and those twenty rows cover the happy path forever. Or production is copied and the names are changed — which puts real customer data on three laptops and in a CI cache, and still does not contain the row that breaks the code, because production never had one.
Across several services it gets worse. Each team builds its own fixtures, so the order the checkout service sends is not quite the order the fulfilment service expects. The integration test that should catch that is written against both teams' assumptions at the same time, so it passes.
What ran
Every table and column with its declared type, nullability, width and precision — and where a CHECK constraint exists, the complete set of legal values. From that come the equivalence classes and boundary values for each column.
Three sources, used in the order the analysis actually resolves them: HTTP routes from the contract shard, which already carry their input, their output and the routines and tables they reach; symbols marked with a data role, for a service that exposes no HTTP at all; and SQL routines, whose parameters, OUTPUT parameters and table touches give all three layers with nothing inferred.
Incoming — what an external system sends. Internal — what the operation works on. Outgoing — what is handed back. These are views, not a partition: an entity reached internally and also returned in a response appears in both, and the overlap is stated rather than left for a reader to discover by adding the counts.
Every value at least once, every pair, every n-way combination, or the full cross-product. The case count and the size of the cross-product it replaces are both reported before anything is generated.
An order always has a customer behind it. Where the schema declares a foreign key it is used; where it does not, a column named after another table's key is derived — and labelled as derived, never as something the schema stated. A polymorphic key that cannot be resolved is left alone and listed.
SQL inserts, JSON fixtures and CSV. Covering records, negative records tagged with the rule each one breaks, and readable filler up to the row target. The whole output is rendered twice and the byte digest compared, so reproducibility is checked rather than asserted.
What it did
Three layers, three kinds of test
The incoming layer is what integration tests need. The internal layer is what service-level tests need. The outgoing layer is what contract tests need. Where the analysis cannot connect an operation to its internals, the endpoint says so on the page — no link is matched by name.
A test file per operation, with the assertion left blank
Each operation with a resolved internal layer gets a test that seeds exactly the rows it touches and calls the operation. The assertion is deliberately absent. What an operation should return is not knowable from a schema, and a generated Assert.Equal(42, result) passes for the wrong reason forever.
What lands
test-data/ the datasets themselves · TEST-DATA-BY-ENDPOINT.html one operation at a time, in its three layers · TEST-DATA-BROWSER.html rows per entity tagged covering, negative or filler, with the constraint behind each · TEST-DATA.md and TEST-DATA-REPORT.html · VALIDATION-REPORT.html
Where it ended
Nothing in the output came from production. Every row exists because a declared type, a CHECK constraint or a threshold in the code put it there, and any value invented to make the data readable is marked as invented — chosen from the column's name and the run's locale, and recorded as generated rather than as something the schema declares. Run it on each service and the outgoing layer of the provider is the incoming layer of its consumer, which is exactly the pair an integration test is supposed to hold together.
What happens next
The reason this beats a fixtures folder is that the rows are derived rather than chosen. A person writing test data writes the cases they have already thought of. A value at exactly the declared width, a null in the one nullable column that matters, a record that breaks precisely one CHECK constraint and nothing else — those are the rows nobody types, and they are the rows that find something.