Clusters and other high performance servers require maintenance. Documented procedures reduce surprises for both enabling scheduled maintenance and emergency work.

Table of contents

Scheduled maintenance and upgrades procedures

Summary

Details

ChemIT notifies cluster lead that maintenance will occur on a specific upcoming date.

Message will state:

Message will be sent to:

Typical work done during maintenance

Sample message:

To: ?
Subject: PI's ClusterName: Date/ time planned down-time.

-----------------------------------------------

To all users of the PI's ClusterName,

On Date/ time, the cluster will be down for planned maintenance for 3 hours.

During this down-time, we intend to:

-----------------------------------------------

 

Emergency work procedures

See also

Communication timeline

1) Something bad happens, which was not scheduled.

2) ChemIT learns of the emergency situation.

3) ChemIT characterizes the problem and develops an initial prognosis.

4) ChemIT notifies group rep (users) of status and prognosis as soon as practicable.

A record (and notes) of Emergency Actions

Cluster or HPC name

Event date
and action

     

Abruna

      

Ananth

      

Collum

      

Hoffmann

      

Lancaster (w/ Crane)

      

Scheraga

      

Widom-Loring

      

Eldor