11 · Case study
An IBM i (AS/400) source exportRPG · CL · DDS · COBOL/400 · SQL · CMD · binder source · JT400Data sheet (PDF)
The system
The problem
An AS/400 combines the language, the database, the screens, the batch scheduler and the packaging into one system, and a source export of it is rarely tidy. Members have been renamed. A shop's own suffixes carry no standard meaning. Source that came out of EBCDIC is frequently not valid UTF-8, and a file decoded carelessly loses the fixed column positions that RPG and DDS depend on. Historical copies sit beside the live member with nothing in the source to say which is which.
Customers also mean different things by moving off the AS/400: a rehost, a refactor in place, a rewrite on a new stack, a package, a data offload, an API layer, or simply capturing what the system does before the people who know it leave. Each of those needs a different part of the picture, and the decision is usually taken before anyone has read the source. This engagement reads it first.
What ran
By extension first, from one configuration file, so adding a member type is a configuration change rather than a code change. Then by content for anything renamed: fixed-format members are recognised by the form indicator in column 6 behind the five-column sequence area, CL by its statement shape, and command source by content, because .cmd is a Windows batch file everywhere else. Content detection is deliberately permissive; the parser makes the final decision, and a member is only recorded as RPG if it parses as RPG. Members whose name marks them as a prior version of a live member are matched to it and flagged.
A file that fails a UTF-8 read is decoded one byte per character. The columns stay where the compiler expects them, which is the difference between a parse and a pile of errors.
RPG in every dialect from RPG III to free-form, including SQLRPGLE; CL; DDS physical, logical, display and printer files; the IBM i bindings of COBOL/400 and ILE C; SQL tables, views, indexes, routines and triggers where a shop has moved off DDS; command definitions and their panel groups; binder source; message files, data areas and data queues; REXX and OCL scripts; the JT400 and XMLSERVICE calls made from Java, Node and Python off the box; and the build metadata in Rules.mk, iproj.json and .ibmi.json. No model is involved at this stage, so nothing in it is inferred.
Program inventory. The call graph, including calls issued from CL, REXX, OCL, ILE C and from off-platform code, merged into the same graph the Workbench uses for every other language. The data model from physical files and SQL tables, with logical files as access paths over their parent. Every record-level read and write, embedded SQL statement and CL file command, located at file and line and classified against the device its target was declared on, so a screen write is never recorded as a database write. The screen inventory with subfile pairs and function keys. Every OVRDBF and OVRPRTF. Submitted jobs, data areas and data queues. Commit and rollback points. Triggers in both forms. The field cross-reference. The message surface. Service program exports and binding directories. The declared object dependency graph and library list.
Every reference to a file, record format or screen that is not in the export is reported by name, with the program and line that reaches for it. A missing object is a finding about the export, and it is listed rather than guessed around.
The same read serves all of them. A rehost needs the object inventory and build dependencies. A refactor in place needs to know which programs are already off the green screen and where the module boundaries are. A rewrite needs the full model, screens, triggers and transaction boundaries included. A package replacement needs the data model and the business rules. A data offload needs the data model and which programs write to each file. API enablement needs the service program exports and the call graph. Knowledge capture needs the field cross-reference and the message surface. The type is an answer the playbook takes, not a different playbook.
What it produced
One model, fifteen sections
Program inventory, call graph, data model, data access, screen inventory, runtime overrides, batch and shared state, transaction boundaries, triggers, field cross-reference, message surface, module boundaries, declared dependencies, missing objects, and the navigation views over the call graph: entry points, reachability, layers, domains, most-changed members, and end-to-end flow and sequence. Each fact carries the member and line it came from. Every section can be read in the Workbench, exported for a report, or handed to another tool.
Impact questions at field level
Which programs reference which column, bounded by the files each program declares. The question is usually asked about a field, and a file-level answer is too coarse to act on.
The mapping for every artifact, on the target you choose
Twenty-one AS/400 artifacts, from an interactive RPG program to a JT400 bridge, each with what it becomes on .NET, on Java, and on any other runtime. The target column used is the one selected when the plan is made, so the decision is taken after the source has been read rather than before.
Where it ended
The target. .NET and Java are the two in normal use for this platform and are named in full; Node and TypeScript, Go, Rust, Python and Scala run on the same mapping with the framework names substituted. DB2 for i maps to SQL Server, Azure SQL, PostgreSQL or Oracle, or stays in place with the application moved first. The data model is produced from DDS and SQL DDL in the same shape either way, so the database decision does not have to be made before the assessment.
What happens next
The rule the playbook follows throughout is that anything derivable from the source is derived, never asked. The model is the finishing layer over the structural read, not the engine underneath it. Two runs over the same export produce the same result, so a figure quoted in an assessment can be reproduced.