When an ERP system stops, the impact is rarely confined to the finance office or IT team. Production orders may disappear from view, stock movements cannot be confirmed, goods cannot be despatched, and planners are left working from incomplete information. The most valuable ERP outage lessons come from treating these incidents as operational failures, not just technical faults.
For manufacturers, ERP and MRP platforms sit at the centre of purchasing, planning, traceability, scheduling and invoicing. A short interruption can create a much longer period of disruption while teams reconcile paper records, re-enter transactions and confirm what has happened on the shop floor. The right response is not simply to restore the server. It is to reduce the chance of recurrence and make recovery predictable when an outage does occur.
1. Define what an ERP outage actually means
An ERP platform can be technically available while still being unusable. Users might reach the login screen but find that reports will not run, barcode scanners cannot post transactions, integrations have failed or performance is so slow that production planning becomes impractical.
That distinction matters when setting service levels and recovery plans. A useful definition of an outage should focus on business capability: can the business release work orders, receive materials, record completed operations, dispatch goods and access the information required for quality and traceability?
Agree these critical activities with operations, production, warehouse and finance leads. Different sites will have different priorities. A make-to-order engineering business may place the greatest emphasis on routing data and job costing, while a high-volume manufacturer may need real-time stock and scanning to remain operational. IT should not make those decisions in isolation.
2. Map every dependency, not just the ERP server
A common lesson from ERP outages is that the application is only one part of the service. The database, virtual host, storage, network, internet connection, identity service, licence server, backup platform and third-party integrations can each create the same visible result: users cannot work.
The dependency list also extends into the factory. Label printers, handheld terminals, shop-floor PCs, shared workstations, warehouse Wi-Fi and interfaces to planning or reporting tools may all depend on the ERP environment. If a production line uses a legacy device or an older operating system, it may need a carefully controlled connection rather than a broad network route to core business systems.
Documenting these relationships gives technical teams a faster route to diagnosis. It also exposes single points of failure before they become an emergency. For example, a database may be protected by backups but still rely on one ageing storage device, or a remote site may have no practical fallback if its sole internet circuit fails.
Include the people dependencies
There is another dependency that is easily missed: access to the people who understand the system. ERP providers, software resellers, database specialists, internal IT teams and managed service providers can all have a role in recovery. If responsibilities are unclear, an incident quickly becomes a sequence of calls asking who owns the next task.
Keep current escalation contacts, support contract details, system credentials held securely and a clear responsibility matrix. During an outage, the priority should be restoring production capability, not establishing which supplier is permitted to investigate.
3. Set recovery targets around production, not convenience
Recovery time objective and recovery point objective are often discussed as technical measures. They are much more useful when expressed in operational terms.
The recovery time objective answers how long the ERP service can be unavailable before the business suffers unacceptable disruption. The recovery point objective answers how much data the business can afford to lose. A four-hour recovery time and a 24-hour data loss window may be tolerable for a low-priority archive, but they are unlikely to suit an ERP system processing stock, goods-in and production activity throughout the day.
The answer depends on how the site works. Some manufacturers can run controlled manual processes for several hours. Others have tightly timed production, traceability requirements or high dispatch volumes that make even a short interruption costly. Be realistic about the trade-off: tighter recovery targets require more investment in resilient infrastructure, replication, backup frequency and tested recovery procedures.
4. Backups only count if they can be restored
Many businesses discover too late that a successful backup notification does not prove the ERP system is recoverable. A backup may be incomplete, encrypted by an attacker, held in the same failure domain as the production system or unable to restore within the required time.
ERP recovery should include the application, database, configuration, custom reports, integration settings and relevant licence information. Restoring a database without the correct application version or interface configuration may leave the business with a technically recovered but operationally incomplete service.
A sensible approach is to maintain protected copies that are separated from the live environment and to test restoration on a planned basis. The test should measure more than whether files appear. Confirm that authorised users can log in, retrieve key records, process representative transactions and operate the integrations that matter to production and dispatch.
Test the awkward scenarios
A power event, failed update and ransomware incident do not behave in the same way. Ransomware may affect servers, user accounts, shared files and backups that remain connected to the network. A failed upgrade may leave a database version mismatch. A storage failure may require a full rebuild of underlying infrastructure.
Testing one scenario is better than none, but rotating through realistic failure modes builds confidence. Record the results, identify delays and assign actions. A recovery plan that has never been tested is an assumption, not a control.
5. Build a manual operating mode before you need it
Paper-based contingency processes are not glamorous, but they can protect output while technical recovery is under way. The aim is not to recreate every ERP function manually. It is to identify the minimum information needed to keep safe, controlled production moving for an agreed period.
This could include printed work orders, controlled stock issue sheets, goods-received logs, dispatch records and a process for capturing completed operations. The process must clearly state who owns each record, where it is stored and how transactions will be reconciled once the system is restored.
Manual workarounds introduce risks of duplication, incorrect stock positions and lost traceability. That is why they need limits. Decide when manual processing starts, which transactions are permitted, how changes are approved and when the business must pause rather than continue on unreliable data. For regulated or quality-sensitive operations, involve the relevant quality lead in designing the process.
6. Separate recovery from incident communication
During an ERP outage, unclear communication creates its own disruption. Production supervisors may make assumptions, sales teams may promise delivery dates without current data and users may repeatedly restart devices or alter records in an attempt to help.
Assign an incident lead who coordinates updates and protects the technical team from constant interruption. Updates should be brief and practical: what is affected, what teams should do now, what workarounds are approved, when the next update will be provided and whether any data-entry restrictions apply.
Communication should also reach third parties where appropriate. If a critical integration partner needs to stop sending data, or if customers need realistic dispatch expectations, that should happen through an agreed business owner. Technical staff should not be left to make commercial decisions in the middle of recovery.
7. Turn ERP outage lessons into funded improvements
The incident review is where an outage becomes useful. Hold it soon after service is stable, while facts are still clear, but avoid turning it into a blame exercise. The question is not who made the mistake. It is which controls, decisions or dependencies allowed a single fault to interrupt operations for so long.
Review the timeline from first symptom to full business recovery. Look for detection gaps, access delays, supplier hand-offs, missing documentation and weaknesses in backup or network design. Then separate immediate corrections from longer-term resilience work.
Immediate actions might include updating contacts, documenting a restart sequence or protecting an overlooked integration. Longer-term work may involve replacing unsupported infrastructure, segmenting shop-floor networks, improving Wi-Fi coverage in warehouse areas, introducing monitored backups or moving a critical workload to a more resilient platform. Prioritise each action by operational impact, risk reduction and realistic cost, rather than trying to fix everything at once.
For businesses with older machinery and specialist applications, modernisation needs particular care. Replacing a legacy device without understanding its production role can create more disruption than the original outage. Segregated networks, controlled jump machines and documented support arrangements can reduce risk while a safe replacement plan is developed.
A dependable ERP environment is built through preparation, tested recovery and clear accountability. The next outage may not be preventable, but it should never be allowed to become a surprise.
