PLAN-NEEDS-TO-LEAVE-PDF

A lot of organisations have an incident response plan.

Far fewer have proved that it works in an OT environment.

And there’s an important difference between the two.

One of the weaknesses I see most often is an incident response plan that is essentially an IT plan with OT added to the title.

The terminology changes.

The systems are different.

There might be a few engineering contacts added to the appendix.

But the response model still assumes cyber teams can contain the problem much as they would in an enterprise environment.

That assumption can become uncomfortable very quickly.

Because in OT, one of the hardest questions during an incident may not be:

How do we isolate the affected system?

It may be:

Can we isolate it without making the situation worse?

Or even:

Should we isolate the plant from the enterprise network to protect it?

That’s when an incident response plan stops being a document and starts becoming an operational capability.

Containment can become an operational decision

Imagine suspicious activity has been identified in the enterprise environment and there is concern it could spread towards OT.

The cybersecurity team quite reasonably starts thinking about containment.

One option is to isolate the plant from the enterprise network.

Close the pathway.

Reduce the exposure.

Protect the plant.

From a cyber perspective, that may sound straightforward.

Then engineering asks:

What happens to the operation if we do that?

Can the plant continue running safely?

Which services disappear?

Which systems degrade?

Does remote support disappear?

Are there dependencies on enterprise authentication, time services, business systems or other shared infrastructure?

And how long can the operation realistically continue in that state?

Suddenly, isolation isn’t simply a firewall decision.

It’s an operational and business continuity decision.

Can you run in island mode?

One question I think every organisation should be able to answer is:

If the enterprise network disappeared tomorrow, could the plant keep operating?

By island mode, I simply mean operating the OT environment with enterprise connectivity removed or heavily restricted.

The important thing isn’t the terminology.

It’s whether you understand the dependencies.

What still works?

What stops?

What degrades?

What becomes manual?

And how long can you sustain it?

A plant may continue operating for an hour without a particular service.

That doesn’t mean it can operate for three days.

And there is a big difference between:

The plant is still running.

and:

The plant can continue running safely and sustainably.

That distinction matters.

Sometimes isolation is there to protect OT

When we talk about isolation during an incident, we often think about containing the compromised system.

Disconnect the host.

Block the communication.

Isolate the segment.

All valid options.

But sometimes the bigger question is the reverse:

How do we protect OT from an incident that started somewhere else?

If ransomware is moving through the enterprise network, can connectivity into OT be restricted quickly?

Who makes that decision?

Does engineering understand what functionality will be lost?

Does cybersecurity understand which connections are genuinely essential?

And has anyone tested whether the plant can still operate afterwards?

This is where segmentation stops being an architecture diagram and becomes part of incident response.

The boundary only really proves its value when you need to use it.

An IT response plan with OT added to the title isn’t enough

An IT playbook might say:

Isolate the endpoint.

Fair enough.

But in OT, what is the endpoint doing?

Is it an engineering workstation?

A historian?

A server supporting multiple production lines?

A jump host used to maintain critical systems?

Can it be disconnected immediately?

What depends on it?

And what happens if you take the action at the wrong time?

The technically correct cyber response can still create an operational problem.

That doesn’t mean you do nothing.

It means the decision needs to consider the operation.

In OT, containment can become an operational decision very quickly.

And if the organisation hasn’t agreed how those decisions get made before the incident, people will be trying to work it out under pressure.

That’s not the ideal time to discover your governance model.

A plan that hasn’t been exercised is still full of assumptions

You can write a very good incident response plan.

Clear diagrams.

Defined roles.

Escalation paths.

Contact lists.

Recovery procedures.

Everything can look perfectly sensible.

Then you run an exercise and somebody asks:

“Who can actually authorise us to isolate the site?”

Or:

“What happens if we lose enterprise connectivity?”

Or:

“How long can production continue like this?”

And the room goes quiet.

That’s useful.

A plan that hasn’t been exercised is still full of assumptions.

Exercises are where you start finding them.

The scenario doesn’t need to be elaborate.

It just needs to force people to make decisions.

Perhaps suspicious activity has been detected in the enterprise environment and there is concern it may reach OT.

Do we isolate the plant?

Do we restrict only certain pathways?

Do we leave connectivity in place while we gather more information?

Who decides?

What operational impact are we prepared to accept?

Then make the situation slightly awkward.

The usual engineering contact is unavailable.

Remote vendor support has gone.

Operations wants to continue production.

Cybersecurity wants faster containment.

Leadership wants a recommendation.

Now the plan has to do some work.

Don’t practise the document. Practise the decisions.

This is the part I think matters most.

An OT incident response exercise shouldn’t simply prove that someone can find the right page in the plan.

It should test whether the organisation can make a safe decision.

Can cybersecurity explain the threat clearly enough for engineering to understand the risk?

Can engineering explain the operational consequence clearly enough for cybersecurity and leadership to understand it?

Does everyone know who has the authority to isolate something?

Can the organisation decide what must remain available?

And can people do that when the obvious answer isn’t available?

That’s where the value is.

You’re not practising the paperwork.

You’re practising how people make decisions together.

Isolation is only half the problem

Suppose the plant has been isolated successfully.

The immediate threat appears contained.

What happens next?

When do you reconnect?

Who decides the enterprise environment is sufficiently trusted?

Does engineering need to validate plant systems first?

Do you restore every connection at once?

Or bring things back gradually?

What does normal look like when connectivity returns?

And what happens if suspicious activity reappears?

Isolation without a recovery plan only solves half the problem.

Recovery is an operational process

A system being technically available does not necessarily mean the operation is ready to use it.

Engineering may need to verify configuration.

The process may need to restart in a particular sequence.

Safety checks may be required.

Operators may need to validate that what they’re seeing matches the physical process.

Vendors or integrators may need to be involved.

The objective isn’t simply:

Get the server running again.

It’s:

Restore the operation safely and with confidence.

That distinction is fundamental in OT.

Final thought

Having an OT incident response plan is a good start.

But the plan itself isn’t the capability.

The capability is whether the organisation can detect something, understand the operational consequence, involve the right people and make a safe decision under pressure.

Sometimes that decision will be to isolate a compromised system.

Sometimes it may be to isolate the plant from everything around it.

Sometimes the safest decision may be to continue operating while more information is gathered.

There isn’t a universal answer.

That’s exactly why the discussion needs to happen before the incident.

Understand what the plant depends on.

Know what can operate independently.

Know what can be safely disconnected.

Practise the decisions.

Then practise the recovery.

Because PDFs aren’t famous for restarting production lines.

Your OT incident response plan needs to leave the PDF.

About the author

Serkan Yusuf is Director of Professional Services at OTIFYD, working with organisations to understand and manage cybersecurity risk across operational technology and industrial environments. OTIFYD helps organisations improve OT cybersecurity and operational resilience through practical, engineering-aware security services.