Many critical-facility incidents occur during planned work rather than unexpected equipment failure. The reason is simple: maintenance and upgrades intentionally change the normal resilient configuration. Change management provides the governance needed to understand that temporary risk before work begins.
What is a change?
A change can include electrical switching, firmware updates, control logic changes, new rack deployment, breaker-setting changes, cooling setpoint adjustment, network changes to BMS or EPMS, replacement of major equipment or even modification of an alarm threshold.
If an activity can alter system behavior, capacity, protection or monitoring, it should be considered for change control.
Classify the risk
Changes should be categorized by operational impact. A low-risk software display correction is not equivalent to isolating one side of a 2N power system. Risk assessment should consider affected systems, redundancy lost, duration, reversibility, human-error potential and customer impact.
Review the MOP
For high-impact work, the MOP should describe the current configuration, intended configuration, prerequisites, sequence, expected indication at each step, hold points, abort criteria and rollback procedure.
The reviewer should ask whether the sequence can create an unexpected common-mode failure or leave the facility in a degraded state.
Confirm capacity before work
When equipment is removed from service, remaining capacity must be sufficient for actual load plus appropriate margin. Verify UPS load, generator availability, cooling capacity, branch-circuit loading and environmental conditions as applicable.
Communication plan
Define who needs to know before, during and after the activity. This can include operations, customers, security, service desk, management and vendors. For high-risk changes, establish a bridge or defined communication channel.
Rollback must be practical
A rollback plan is useful only if it can actually be executed. Required tools, spare parts, access, software backups and qualified personnel should be available before the change begins.
Freeze conditions
Some changes should not proceed during adverse weather, major business events, reduced staffing, another concurrent maintenance activity or when redundancy is already degraded. The go/no-go decision should be made immediately before starting work.
Configuration control
After a successful change, update drawings, asset records, settings, labels, monitoring points and procedures. A technically successful change that leaves documentation outdated creates future risk.
Post-change verification
- Confirm all equipment returned to the intended state.
- Verify alarms and temporary bypasses are cleared.
- Check redundancy and capacity.
- Review monitoring trends for abnormal behavior.
- Update documentation and asset records.
- Close the change only after final verification.
Lessons learned
If the work deviated from plan, record what happened. Near misses and unexpected responses should improve future procedures even when no outage occurred.
Key takeaway
Change management is not administrative paperwork. It is the process that converts planned technical work into controlled risk. The most important decisions often happen before the first switch is operated: understanding the system state, verifying capacity, validating rollback and deciding whether conditions are safe to proceed.
References and Further Reading
- ISO/IEC TS 22237-7:2018, data center management and operational information.
- ISO 55001:2024, Asset management system requirements.