Production network
Convenience
Finding an unstable network fault without stopping production
We found the cause, replaced the equipment hiding it and introduced alerts that people could act on.

Short network failures affected production, but the pattern was too inconsistent for a quick diagnosis.
Measure before replacing
The visible symptom appeared in several places. We collected enough evidence to separate the network fault from application and workstation issues.
Change one dependency at a time
The faulty equipment was replaced in a planned window. We then tested the production flow, not only the connection status.
Reduce alert noise
Notifications now represent an incident and its recovery instead of sending repeated messages for the same event.

