
Legacy modernisation with coding agents
Year:
2026
Service:
AI-assisted Engineering
Industry:
SaaS
Team:
4 engineers, 14 weeks
Reference scenario: a UK SaaS vendor moves the legacy billing core of its product to .NET 10 with coding agents. Tests capture the old behaviour first, agents migrate feature by feature, and engineers review every change and lead the complex features.
Introduction
In this scenario a UK-based SaaS company sells a billing and field-service platform to telecom operators and utilities in the UK and the EU. The web layer is modern. The rating and invoicing core is not: fifteen-year-old VB.NET on .NET Framework, business rules buried in stored procedures, almost no automated tests and two engineers who understand it. A manual rewrite has been estimated twice and postponed twice.
The head of engineering reopens the question because coding agents have changed the sums. A published industrial case study, in which a coding agent migrated 12 features of an ERP system from Visual Basic 6 to C# on .NET 10, reports 92% functional equivalence for low-complexity features and 47% for high-complexity ones. That is the honest picture: agents can carry most of the simple work, and the hard part belongs to engineers.

Challenge
No definition of correct. The old core has no test suite, so nobody can say what equivalent behaviour means. Migrating without that definition would faithfully carry over bugs and dead code, or quietly change invoices.
Uneven agent performance. Research cited in DORA's 2026 report on return on investment finds productivity gains of 35 to 40% on simple greenfield tasks and 10% or less on complex legacy code. METR's May 2026 report puts public frontier models at a task horizon of about 12 hours at 50% success, but only about 1.5 hours at 80%. Long autonomous runs on a billing engine are out of the question; the work has to be cut into short steps that can be checked.
Agents close to production data. Customer contracts forbid passing billing data to external tools, and the head of security will not give an agent credentials that reach production.
No big bang. Operators invoice their own customers every month. The cut-over has to be gradual and reversible.
Solution
The order of work matters more than the choice of agent.
Tests first, agents second. We record real request and response traces from the old core, anonymise them and turn them into characterisation tests and golden-master suites. That produces an executable definition of equivalence, and exposes dead code that is retired instead of migrated.
Triage by complexity. Every feature is scored. Simple ones go to agents working to a written specification; complex ones are led by engineers, with agents assisting.
Review independent of the author. Every agent-written change arrives as a pull request, is checked by a verifier agent that did not write it, and is approved by an engineer.
What gets built:
A coding agent workflow (Claude Code) with repository-level instructions and reusable skills for the migration patterns.
MCP servers that give agents read-only access to legacy documentation, database schemas, the issue tracker and CI results.
Sandboxed runners with short-lived, narrowly scoped tokens; synthetic and anonymised data only; no route to production.
An equivalence harness in CI that replays the golden-master suite against the old and new implementations and reports a result per feature.
A strangler facade with feature flags: each migrated feature takes traffic gradually and can be switched off.
A dashboard with DORA delivery metrics, equivalence per feature, review time and token cost.
Stack: .NET 10, C#, Claude Code, MCP, CI with sandboxed runners, feature flags.

Result
Acceptance criteria and targets in the scenario:
Cut-over criterion: a feature takes production traffic only when it passes 100% of its golden-master tests and a period of shadow traffic with no invoice differences.
First-pass target: at least 90% equivalence for low-complexity features before human fixes, in line with the published case study; no target is set for complex features, which engineers lead.
Delivery targets: lead time for changes in the migrated modules shortens from release to release, and the change failure rate does not rise. Both are compared with a baseline measured in week one.
Cost model: token cost is logged per feature; the planning reference is the study's figures of about USD 1.66 for a low-complexity feature and USD 10.28 for a high-complexity one.
In the scenario the vendor finishes with the first modules in production on .NET 10, a test suite the old system never had, an agent workflow its own engineers run and a measured map of what remains.

