Individual equipment testing can prove that a UPS, generator, chiller or ATS works on its own. It cannot prove that the data center will remain stable when several systems must respond to the same event. Integrated Systems Testing (IST) addresses this gap by testing complete operational scenarios across system boundaries.
Start with a defined objective
Every IST scenario should answer a specific resilience question. Examples include: Can the facility survive loss of utility power? Can cooling recover after generator transfer? Does one failed generator leave sufficient capacity? Does a control-network failure create unsafe equipment behavior?
Build the scenario from the design
The test should reflect actual one-line diagrams, mechanical architecture, redundancy philosophy and control sequences. Artificial scenarios that cannot occur in the real system may have limited value unless they are intentionally testing a protective boundary.
Failure injection must be controlled
IST deliberately creates abnormal conditions. The test plan should therefore define prerequisites, risk controls, personnel, communication channels, abort criteria and rollback steps before the first failure is introduced.
The safety of people and critical equipment takes priority over completing a test script.
Utility-loss scenario
A representative power scenario can include loss of normal supply, UPS transition to stored energy, generator start, synchronization or ATS transfer, UPS acceptance of generator power, cooling restoration and eventual return to normal utility.
Operators should record actual timing, alarms and transient behavior instead of simply marking the test as pass.
Cooling-failure scenarios
Cooling IST can include loss of one chiller, pump, CRAH group, control sensor or electrical source. The test should determine whether standby capacity starts, whether rack inlet temperatures remain acceptable and whether the facility returns to a stable condition.
Redundancy scenarios
Testing should examine the facility in degraded states as well as normal full-redundancy conditions. If one UPS, generator or cooling component is already under maintenance, what happens when another fault occurs?
Monitoring is part of the test
BMS, EPMS and DCIM should capture the event accurately. Incorrect timestamps, missing alarms or stale data can make a technically successful infrastructure response difficult to operate during a real incident.
Operations must participate
IST is also an operational-readiness test. The operations team should practice EOPs, communication, escalation and decision-making. A system may respond correctly while operators remain unprepared to interpret what they see.
Recovery is part of the scenario
- Verify equipment returns to normal mode.
- Confirm temporary bypasses are removed.
- Check redundancy is restored.
- Verify batteries or fuel systems recover appropriately.
- Confirm alarms clear correctly.
- Review monitoring trends after stabilization.
Key takeaway
IST proves interfaces, timing and recovery under realistic stress. Its value is not in creating dramatic failures, but in producing evidence that the entire facility responds predictably when normal conditions disappear. A successful IST demonstrates both technical resilience and operational readiness.
References and Further Reading
- ASHRAE Guideline 0-2019, The Commissioning Process.
- ANSI/ASHRAE/IES Standard 202-2024.
- ISO/IEC 22237-1:2021, Data centre facilities and infrastructures — General concepts.