CogniDev Workbench
Open a repository and the whole system is read first: the layers, the flows, and where a change will land. The workbench then proposes the work, writes the plan, builds it task by task and produces the evidence. Every gate is yours to approve.
Runs on your machine. Works with any model. Shared across your team.
The same playbooks, analysis and reports on all three.
01 · Capabilities
Each area below is covered by a playbook: a defined process, with the stacks it has been proven on.
Start a new service with its architecture, tests, CI and delivery already in place.
Next.js · Java · .NET · Go · Rust · Scala · Node · Python · Terraform
Begin from an industry model rather than an empty repository. The capabilities, entities and business flows are already defined.
Capital Markets — 50 capabilities across 7 desk areas · custom blueprints on any runtime
Map an undocumented system: layers, flows, use cases, routes, data model and call graph, each traced to the files it comes from.
Java · .NET · Node / TypeScript · Python · Go · Rust · Scala · COBOL · RPG · CL · DDS · C · C++ · SQL
Add to an existing codebase following its own layers and conventions, with tests and wiring included.
Next.js · Java · .NET · Go · Rust · Scala · Terraform
Bring an AI-generated prototype up to a standard a team can maintain: structure, code quality, tests, lint, CI, documentation and deployment. Issues are repaired and re-checked, not only reported.
Ten checks, each one fixed and verified before the next
Move framework and language versions forward on the same system, with the build passing at every step.
Java → current LTS · .NET → latest · AngularJS → Angular 22 · Scala · dependency uplift
Move a system onto a supported stack. Behaviour is carried across and verified as equivalent.
COBOL → Java · COBOL → .NET · IBM i (RPG, CL, DDS) → .NET or Java · VBScript / classic ASP → .NET 10 · AngularJS → React
Split a monolith along boundaries derived from the code: services, contracts, data ownership and the transactions between them.
Java · .NET · Node / TypeScript · Python
Move a schema between dialects, move or transform the data behind it, and bring logic held in stored procedures into application code.
Dialect → dialect · data movement · stored procedures → .NET · schema and data-access mapping
Expose an API over an existing domain, or bring the APIs you already have under a single governed contract.
API-first for .NET · expose over an existing domain · API Studio · contract governance
Build the test coverage a system is missing, and verify a migration by comparing the new system's behaviour against the original.
Selenium → Playwright · tests → Gherkin · equivalence checks · test data · coverage gates
Score every repository and enforce thresholds in CI. Findings are ranked by reachability, so the list stays actionable.
Secrets · CVEs / SBOM · SAST · OWASP Top 10 · scorecards · SOC 2 · HIPAA · PCI-DSS · GDPR · ISO 27001
Each run produces a branch, one commit per verified task, and a report committed alongside your code. See the full library →
02 · Method
Every capability above runs as a playbook: a versioned process that defines the steps, the approval gates, the tools it runs and the standards it enforces. The same job runs the same way from the desktop app, the command line or CI.
Playbook structure
Every run, the same six gates
The structural map, and a plain-language statement of what the codebase is and does.
The options the code supports, each with the evidence that qualified it.
A file-level plan, written down before anything changes.
You read the plan and approve it. Nothing is written until you do.
Task by task. Each one compiles and typechecks before it lands, one commit per passing task.
A report of files, commits, tasks and checks. The work lands as a pull request and is never auto-merged.
Delivery and updates
Each playbook is authored in advance: the ordered steps, the approval gates, the tools it may run, the standards it enforces and the questions it is permitted to ask. None of it is decided at run time.
Your architecture rules, reference implementations and coding standards are placed alongside a playbook, and the run is held to them. The playbook itself is not forked or rewritten.
Opening a folder reads the code before anything is offered. Only the playbooks the evidence supports are listed, each with the reason it qualified. Nothing runs until one is selected.
Playbooks are versioned and signed, and are delivered over the air. A correction authored once reaches every installation without a reinstall, and each run records the version it used.
Where the token cost goes
How much this saves depends on the repository, so the workbench meters it rather than quoting a figure: every call reports its tokens and its cost against the step that made it.
Desktop app, command line, Docker, or called as a skill by a coding agent — the same playbook, the same output. Models are your own: Anthropic, Bedrock, Copilot, OpenRouter, z.ai, Kimi, or any OpenAI-compatible endpoint, including Ollama and vLLM on your own hardware. Nothing leaves the machine you ran it on.
03 · Comparison
Assistants and agent skills both put a model in charge of the whole run, so the same request can produce a different result each time. A playbook keeps the process fixed and calls the model where judgement is actually needed.
| CogniDev WorkbenchA playbook | AI coding assistantsCursor · Copilot · Windsurf | Agent skillsClaude Skills · MCP skills | |
|---|---|---|---|
| Unit of work | A playbook — a versioned process with ordered steps and approval gates | A prompt | A set of instructions the model reads and interprets |
| What drives the run | The playbook. The model is called at the steps that need judgement, inside a fixed process | The model, end to end | The model, end to end |
| Same input, same result | Yes. The map, the options and the plan are derived from the code | No | No |
| Architecture | Proven industry architectures, selected per project, with the trade-offs recorded | Whatever the prompt describes | Whatever the instructions describe |
| Scaffolding and code generation | Dedicated parsers and generators, versioned with the playbook | Generated text | Generated text |
| Token cost | Paid once for understanding, then per judgement call, metered per step | Re-paid on every prompt, by every developer | Re-paid on every invocation |
| Codebase understanding | Parsed and written to disk, updated incrementally, shared across the team | Retrieved per prompt, discarded after | Retrieved per invocation, discarded after |
| Your standards and references | Injected alongside the playbook and enforced at every gate | Pasted into the prompt | Written into the instructions |
| Long-running work | Resumable, verified per task, one commit per passing task | One conversation at a time | One session at a time |
| What a reviewer receives | The plan, per-task commits, a verification log and a report | A chat transcript | A chat transcript |
| Unattended runs | Command line and Docker — the same playbook, the same output | Interactive | Depends on the host |
The difference is where the model sits: inside a defined process, at the steps that need judgement, rather than driving the run from end to end.
Get started
Tell us your stack and what you need done, and we will point you at the right playbook.