At 06:12, a production supervisor cannot access the MRP system. Labels will not print, stock movements cannot be recorded and the morning shift is left waiting for instructions. Whether the cause is ransomware, a failed server, flood damage or an accidental configuration change, a manufacturing disaster recovery plan determines whether this becomes a short disruption or a costly stoppage.
For manufacturers, recovery is not simply about getting files back. It is about restoring the order of operations safely: communications, identity access, ERP or MRP, production scheduling, warehouse systems, machine connectivity and the data needed to prove what has been made. A plan that only covers office IT can leave the factory exposed when it matters most.
What a manufacturing disaster recovery plan must protect
A disaster recovery plan should identify every system that affects output, not just the systems owned by the IT department. That includes servers and cloud platforms, but also shared shop-floor PCs, barcode scanners, label printers, warehouse Wi-Fi, VPN access, engineering workstations and interfaces between machinery and business applications.
The difficult part is that each asset has a different recovery requirement. An email outage may be inconvenient for a few hours. Loss of the ERP database during a production run may affect purchasing, traceability, despatch and invoicing. A compromised machine-side PC may require a controlled response involving the equipment supplier before it can be restarted safely.
This is why recovery planning needs operational input. Production, engineering, warehouse, finance and IT should agree which systems are critical, who can authorise recovery decisions and what an acceptable period of manual working actually looks like.
7 tests for your manufacturing disaster recovery plan
1. Can you name the systems that stop production?
Start with a practical impact assessment. Ask what prevents the next job from being released, made, checked, labelled or despatched. The answer will usually be a chain of systems rather than one application.
Document the ERP or MRP platform, file shares containing drawings and work instructions, scheduling tools, quality records, warehouse devices, remote access, internet connectivity and key supplier-managed systems. Include dependencies such as Active Directory, DNS, network switches and Wi-Fi. A production application cannot be restored if users cannot authenticate or the shop floor cannot reach it.
Give each system a named business owner. IT can explain technical dependencies, but the business owner should set the recovery priority based on operational impact.
2. Do your recovery times reflect the cost of downtime?
Two measures should shape the plan: recovery time objective and recovery point objective. The recovery time objective is how quickly a system must be available again. The recovery point objective is how much data loss the business can tolerate.
For example, a daily backup may be suitable for a file archive but unacceptable for live order processing if a full day of transactions would need to be recreated. Equally, an expensive high-availability arrangement may not be justified for a system used only occasionally. The right answer depends on the effect of downtime, the availability of manual workarounds and the cost of restoring data inaccurately.
Set targets in plain terms. Rather than saying a system is “high priority”, state that it must be available within four hours and that no more than one hour of data may be lost. Those targets make backup, cloud and support decisions easier to assess.
3. Have you tested that backups can actually be restored?
A backup report showing a green tick is useful, but it does not prove that the application will run after restoration. Backups can be incomplete, encrypted by an attacker, stored with the wrong permissions or unable to restore a specific database and its supporting files.
Test recovery at least quarterly for critical systems. Restore a representative copy in a safe environment and check that users can log in, records are current and the application performs as expected. For ERP and MRP platforms, involve the software provider where necessary, particularly if licensing, database versions or custom integrations are involved.
Keep at least one protected copy that cannot be altered by a compromised administrator account. This is a core defence against ransomware, which often targets backup repositories before demanding payment.
4. Can the factory operate safely while systems are unavailable?
Not every incident can be fixed immediately. Your plan should set out a controlled manual process for the first hours of an outage, including how work orders are issued, how materials are booked out, how quality checks are recorded and how finished goods are tracked.
Manual working creates risk. If paper records are later entered twice, traceability can be damaged and stock figures can become unreliable. Prepare numbered forms, clear ownership and a reconciliation process for when systems return. The aim is not to make manual work convenient. It is to keep production moving without creating a second problem to solve afterwards.
For machinery connected to older operating systems, avoid improvised changes during an incident. A machine may have been validated with a particular configuration, and restarting or patching it without the correct process can introduce safety, warranty or compliance concerns.
5. Are shop-floor networks separated from office systems?
Network segregation limits the spread of a cyber incident and makes recovery more controlled. Office users, guest devices, servers, production equipment and building services should not all sit on the same unrestricted network.
Older machinery may be unable to support modern security controls. That does not mean it must be ignored. A segregated network, tightly controlled firewall rules, monitored remote access and a managed jump machine can reduce exposure while preserving a stable machine environment.
Recovery planning should include network diagrams, firewall configuration backups and a route for safely reconnecting production assets. If a ransomware incident affects office devices, the response may be to isolate affected areas quickly while maintaining only the essential, approved connections to operational technology.
6. Does everyone know who makes decisions?
During an outage, uncertainty wastes time. The plan should name the incident lead, technical recovery lead, production decision-maker, communications owner and external contacts for software, machinery and telecoms suppliers.
Write down escalation routes and out-of-hours contact details. Clarify who can approve a shutdown, who can authorise recovery from backup and who speaks to customers if delivery dates are affected. This is particularly valuable where internal IT, a managed IT provider and several specialist vendors all support different parts of the environment.
A simple incident log also matters. Record what happened, when systems were isolated, which actions were taken and who approved them. It supports a calmer handover between shifts and provides evidence for later review.
7. Have you rehearsed a realistic failure?
A printed plan is not proof of preparedness. Run a tabletop exercise based on a credible scenario, such as a ransomware alert on a shared production PC, loss of the ERP server or an internet outage affecting cloud applications and remote support.
Walk through the first hour, the first shift and the first day. Test whether contact numbers work, whether decision-makers know their role and whether recovery priorities still make sense. Then test a technical element, such as restoring a critical database or switching to a backup internet connection.
Treat findings as improvement work, not blame. Plans become outdated when new machinery, software integrations, users or suppliers are introduced. Review the plan after major changes and after every real incident, even if the disruption was brief.
Recovery planning is also a security decision
Many disasters begin as security incidents. Weak passwords, unpatched remote access, shared accounts and unrestricted network access can turn one compromised device into a wider operational outage. Prevention reduces the chance of recovery being needed, while recovery limits the damage when prevention fails.
This is where Cyber Essentials controls, multi-factor authentication, managed patching, endpoint protection and access reviews have direct production value. They are not separate compliance exercises. They help protect the systems that issue work, hold drawings, manage stock and connect teams across the factory and office.
A specialist manufacturing IT partner can also help define the boundary between supported business systems and legacy equipment that needs a more cautious approach. The objective is to improve security without making untested changes to machinery that production depends on.
Make the plan usable under pressure
The best plan is short enough to use at 06:12 on a difficult morning. Keep the first-response actions, contact list, system priorities and recovery decision tree in a secure location that remains available if normal systems are down. Store an offline or protected copy as well as a controlled digital version.
Then give the plan an owner and a review date. Recovery capability is built through tested backups, clear responsibilities and decisions made before the pressure arrives. When the next failure occurs, your team should not be deciding what matters most. They should already know how to protect output and restore it safely.
