Clusters and other high performance servers require maintenance. Documented procedures reduce surprises for both enabling scheduled maintenance and emergency work. |
ChemIT notifies cluster lead that maintenance will occur on a specific upcoming date.
Message will state:
Message will be sent to:
(Point to Lulu's current, active checklist! Nov 2015)
Update BIOS (and why it's done...)
Update OS
Update...
Test UPS
Test backups
Test...
To: ?
Subject: PI's ClusterName: Date/ time planned down-time.
-----------------------------------------------
To all users of the PI's ClusterName,
On Date/ time, the cluster will be down for planned maintenance for 3 hours.
During this down-time, we intend to:
-----------------------------------------------
1) ChemIT notifies group rep. of planned date.
2) Group rep. confirms there is no better date (or negotiates a better date, with ChemIT staff).
3) Group rep. notifies all users of cluster, using message crafted by ChemIT.
4) The work day before the shut down, ChemIT sends a reminder.
5) When cluster is shutdown, ChemIT sends a statement to that affect.
6) ChemIT sends a status report if cluster not up when expected, providing new time estimate.
7) ChemIT sends a report when the server is again available.
1) Something bad happens, which was not scheduled.
2) ChemIT learns of the emergency situation.
3) ChemIT characterizes the problem and develops an initial prognosis.
4) ChemIT notifies group rep (users) of status and prognosis as soon as practicable.
Cluster or HPC name | Event date | |||||
|---|---|---|---|---|---|---|
Eldor |