The operational problem
Industrial recovery has a milestone that conventional IT recovery often does not: a system can be technically available while the physical process is still not safe enough, trustworthy enough or understood well enough to resume. A virtual machine may boot, a PLC may answer and a backup may restore cleanly, yet the plant still needs confidence in recipes, logic, remote-access paths, engineering workstations and privileged identities.
That distinction is consistent with NIST SP 800-82 Rev. 3, which treats OT security as inseparable from reliability and safety constraints. Recovery therefore cannot be reduced to server availability. It needs evidence that the restored technical state matches an acceptable operational state.
A defensible restart decision depends on evidence that survives the compromised domain and can be interpreted consistently by IT and OT.
Do independent cyber and process records converge on one defensible plant state?
What trustworthy restart evidence requires
The first rule is independence. If the hypervisor, identity plane or engineering workstation sits inside the suspected compromise boundary, relying exclusively on logs produced by that same domain creates circular assurance. Restart-critical evidence should be tamper-evident and, where practical, collected or replicated out of band.
The second rule is provenance. A recently approved baseline is not automatically known-good. If the possible attacker dwell time predates the baseline, that baseline becomes part of the investigation. The useful comparison points are independent PLC logic, signed firmware, validated recipes, remote-session records and privileged-account history. ISA/IEC 62443 provides the broader lifecycle and system-security context, while NIS2 reinforces the governance expectation around resilience and incident handling.
The third rule is a shared IT/OT decision language. At 03:00, a SOC analyst may interpret a connection as persistence while an operator sees a normal industrial cycle. Pre-agreed go/no-go criteria, evidence owners and restart authority reduce the need to renegotiate acceptable risk while production is waiting.
- Define which PLC logic, recipes, firmware, accounts and remote sessions are restart-critical.
- Protect evidence sources from the same administrative domain they are intended to validate.
- Assess whether baselines predate the possible attacker dwell time.
- Agree IT/OT interpretation rules for common industrial behaviours before an outage.
- Name the individual or forum with authority to accept residual restart risk.
