Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

P/T: Default answer is "no action". This is here in case Is there a Project or Ticket has an impact on this topic, which is to be done awaiting implementation during the maintenance work.?

D: Discretion or decision required. May very well be, "nothing this time around". And just because something can be done, doesn't necessarily mean it should be done: risk/ benefit.

Sequence of actions taken during maintenance

...

Actionnotes
Schedule / notify representative / users 
Evaluate hard drives 
Verify backups

...

 
Disable outside access 

...

Delete jobs

...

 

...

Shutdown headnode & nodes 

...

Synology update

...

...

Reboot switches

...

 
Boot synology 

...

Boot headnode

    o iSCSI config- (on startup) 5 sec ping for timeout, will change to 10 & 15

· Test UPS shutdown error

· Reboot anything that needs

· Boot nodes

· Enable access

· Send email

    • UPS config - 10 seconds poll,  

...

 
Verify drives with fsckOnly do every ~6 months. Time consuming.
Ex: Abruna's cluster (small): 1-2 hours
Test UPS and its notificationsDoes it work as expected? How reasonably test?
Reboot anything that needs 
Boot nodes 
Enable access 
Send email