...
P/T: Default answer is "no action". This is here in case Is there a Project or Ticket has an impact on this topic, which is to be done awaiting implementation during the maintenance work.?
D: Discretion or decision required. May very well be, "nothing this time around". And just because something can be done, doesn't necessarily mean it should be done: risk/ benefit.
Sequence of actions taken during maintenance
...
| Action | notes |
|---|---|
| Schedule / notify representative / users | |
| Evaluate hard drives | |
| Verify backups |
...
| Disable outside access |
...
| Delete jobs |
...
...
| Shutdown headnode & nodes |
...
| Synology update |
...
...
| Reboot switches |
...
| Boot synology |
...
| Boot headnode |
o iSCSI config- (on startup) 5 sec ping for timeout, will change to 10 & 15
· Test UPS shutdown error
· Reboot anything that needs
· Boot nodes
· Enable access
· Send email
• UPS config - 10 seconds poll,
...
| Verify drives with fsck | Only do every ~6 months. Time consuming. Ex: Abruna's cluster (small): 1-2 hours |
| Test UPS and its notifications | Does it work as expected? How reasonably test? |
| Reboot anything that needs | |
| Boot nodes | |
| Enable access | |
| Send email |