On monitoring


Monitoring Job 1: Daily

Frequency: Every 1 day

Send via e-list: DNS-DIAPER-REPORT-L@cornell.edu

Content: Daily data entry count of the last 30 days


Monitoring Job 2: Hourly

Frequency: Every 1 hour

Step 1: Init

alert_count = 0

Step 2: Run all rules

TargetData sourceValueExpected (if)
& Otherwise
Action (then)Personnel
DatabaseBioHPC: test.sampleCount of entries filtered by:
(date_time in last 48 hours)
>= 1// Do nothing



== 0 (or timeout)alert_count + 1All
Backend mobile APIsmobile test server
/api/monitoring
{
  "api_env": "test",
  "connected_db": "test",
  "pipeline": "mobile"
}
Match all keys/values// Do nothing



else (or timeout)alert_count + 1Chengyong, Vivian

mobile prod server
/api/monitoring
{
  "api_env": "production",
  "connected_db": "production",
  "pipeline": "mobile"
}
Match all keys/values// Do nothing



else (or timeout)alert_count + 1Chengyong, Vivian
Backend dashboard APIsdashboard test server
/env
{
  "api_env": "TEST",
  "connected_db": "TEST",
  "pipeline": "dashboard"
}
Match all keys/values// Do nothing



else (or timeout)alert_count + 1Weihang

dashboard prod server
/env
{
  "api_env": "PROD",
  "connected_db": "PROD",
  "pipeline": "dashboard"
}
Match all keys/values// Do nothing



else (or timeout)alert_count + 1Weihang

Step 3: Determine status & Send emails

if (alert_count > 0) {
	put all error info (expected v actual) in an email
	send the email to DNS-DIAPER-ALERT-L@cornell.edu
}
else if (current time between 08:50am and 09:10am) {
	send an email "All status OK" to DNS-DIAPER-REPORT-L@cornell.edu
}
// else { do nothing; will run again at next hour; }