Network monitoring should answer three questions: is the service available, is performance within acceptable limits, and what evidence exists when something changes? Device-up status alone cannot provide those answers.
Monitor the full service path
Collect health and availability data from switches, routers, firewalls, load balancers, carrier interfaces, management systems and critical network services. Monitoring should include both A and B paths where they exist.
Use interface and traffic telemetry
Important indicators include utilization, errors, discards, link-state changes, optical receive/transmit levels where supported, queue behavior and packet loss. Flow records and application measurements can help identify abnormal traffic patterns and congestion.
Monitor control-plane health
Routing adjacency state, route counts, convergence events and protocol changes can reveal problems not visible from physical link status. Repeated flapping should be investigated even when redundancy prevents an immediate outage.
Synchronize time
Accurate timestamps are essential for correlating network logs with server, security, BMS, UPS and application events. Time synchronization should therefore be treated as a critical supporting service and monitored for failure.
Design useful alarms
Alarm thresholds should reflect operational risk. Excessive low-value alarms train teams to ignore notifications, while missing alarms delay response. Alerts should identify the affected device or service, severity, time and relevant dependency.
Retain incident evidence
Configuration history, event logs, interface counters, routing changes and monitoring trends should be retained for a period appropriate to operational and compliance requirements. After incidents, evidence should support a timeline, root-cause analysis and corrective actions.
References and further reading
- ISO/IEC TS 22237-7:2018, Management and operational information.
- ANSI/TIA-942-C, Telecommunications Infrastructure Standard for Data Centers.
- IETF operational RFCs for the routing protocols implemented in the environment.