cutover-community
Blog
October 1, 2026

Every DR Test Makes the Next One Smarter: Inside Cutover's AI Learning Cycle for Resilience

Wave one of a data center disaster recovery (DR) test skips the failover rehearsal. So does wave two. So does wave three. Nobody decided that. It just happened and the lesson got buried in a yellow highlighted cell on a spreadsheet that nobody will re-open at 10pm on a Saturday, mid-recovery. That's not a training problem. That's an architecture problem. Most disaster recovery tooling has no memory. Every run starts from zero, no matter how many times you've run it before.

That's the gap self-improving disaster recovery is built to close. Not automation that runs a fixed plan faster.  Rather, its intelligence that changes/updates the plan itself, based on what actually happened last time. See what that looks like in practice.

The old model: static IT disaster recovery plans, static lessons

Most DR and resilience programs run on a version of the same broken loop. Build a plan. Execute it. Write a postmortem nobody reads. Copy the spreadsheet forward to the next wave, mistakes and all. The data about what went wrong is there, it's just inert. It sits in a cell, a Slack thread, a slide nobody opens again. The team that skipped the rollback rehearsal in wave one skips it again in wave two, and again in wave three, because the plan itself never learned anything.

That's not a resourcing problem. It's a structural one: plans don't remember, and people forget under pressure.

The shift: intelligence, under governance

Turning every DR test and every incident into a durable learning advantage requires two things most platforms don't have together: AI that can act inside a live operation, and hard governance over what that AI is allowed to do without a human saying yes. One without the other doesn't work. Full autonomy in a live failover is a liability. Full manual review of every AI suggestion defeats the point.

Here's the loop, in practice, inside Cutover Recover and Cutover AI:

  1. AI Create - build the runbook. Feed it your legacy documents such as  spreadsheets, Word docs, PDFs, etc and one paragraph of intent. In under a minute, you get a complete runbook: tasks, dependencies, streams, a go/no-go gate. It even catches the mistake buried in the yellow cell such as a DNS propagation window planned for 10 minutes that should be 30.
  2. Agentic tasks - execute with a memory. These sit inside the runbook itself. When execution reaches one, an agent, whether it is Cutover's built-in agent or your own already-approved internal agent, takes over. No human has to remember to ask it to act; it acts automatically when the runbook reaches it, under the same approvals your organization already set.
  3. Evaluate - before the decision, not after. Ahead of a go/no-go gate, an agentic readiness check reviews every prior wave automatically: what overran, what got skipped, and why plus citing each wave by name. The Major Incident Manager walks into the decision with a briefing, not a guess. The final call stays human.
  4. Improve - propose, don't impose. After execution, a second agentic task compares planned versus actual durations across every wave, calculates medians, and proposes updates to the master template. Nothing changes the plan without a person approving it. Once approved, the change gets applied and logged, and a clean wave report lands automatically. No one spends hours building a postmortem deck.
  5. “Rinse and Repeat” - the next run starts smarter. The rollback rehearsal that got skipped three times in a row is now scheduled automatically, with a real-world duration attached. That's the recursion. Every test and every incident becomes training data for the next one.

Why governance is the whole point, not a constraint on it

Generic AI dropped into a live incident or a live failover is a bet, not a plan. Your enterprise IT estate is a snowflake and generic AI trained on the public internet doesn't know your architecture, your legacy code, or your internal acronyms. Foundation models know what's written. Cutover captures what's done: real operational sequences, not logs.

That's why every agentic task in the loop proposes, and a human approves. Trust is earned incrementally.  It is assistive first, then supervised, then autonomous only where the risk profile allows it. The recursion is real, but it's bounded: an agent can act automatically inside its granted authority, and it can recommend a change to that authority, but it can never expand its own authority. That distinction is what makes it safe to run in production, and it's what makes the learning durable instead of reckless.

The compounding advantage

Every failover, every DR test, every incident you run without this loop is a lesson your team has to relearn the hard way; again, under pressure, at 2am. Run it with recursive intelligence and governance built in, and the operational knowledge compounds: fewer skipped steps, tighter time estimates, templates that reflect what your organization actually does instead of what someone thought it would do a year ago. That accumulated execution history, not the model underneath it, is the advantage that doesn't get copied.

Frequently Asked Questions

What makes disaster recovery "self-improving" instead of just automated?
Automation executes a fixed plan faster. Self-improving disaster recovery changes the plan itself, based on what actually happened in prior runs, with a human approving every change before it goes live.

Does the AI make decisions during a live cutover without approval?
No. Agentic tasks act automatically when execution reaches them, but only within pre-approved guardrails. Anything that changes the plan such as a template update, a new task, a policy change requires human sign-off.

How is this different from a standard post-incident review?
A manual postmortem is a one-time document someone has to write and someone else has to remember to read. Cutover's agentic review runs automatically before every go/no-go decision and after every wave, citing prior runs by name, and it feeds directly back into the template so no review, no update required to make it happen.

What data does the AI use to make these recommendations?
Your own execution history: planned versus actual task durations, what was skipped, what overran, and why.  All of that is captured automatically as a byproduct of running the Cutover runbook, not reconstructed after the fact.

Is this only for disaster recovery?
No. The same loop applies to migrations, cyber recovery, and release management, basically any recurring, high-stakes technical operation executed from a runbook.

See it run

See how Cutover turns every test and incident into a durable learning advantage: book a demo.

Kieran Gutteridge
IT disaster recovery
AI
Latest blog posts