Closed Bug 1116990 Opened 11 years ago Closed 11 years ago

badg.us & discourse Pingdom alerts, RCA

Categories

(Participation Infrastructure :: MCWS, task)

x86
macOS
task
Not set
normal

Tracking

(Not tracked)

RESOLVED WONTFIX

People

(Reporter: mrz, Unassigned)

Details

Pingdom noted two alerts last night (2014/12/31) US/Pacific, hours apart. Both recovered quickly. discourse.mozilla-community.org is DOWN since 2014-12-31 21:17:09 GMT +0000 discourse.mozilla-community.org is UP since 2014-12-31 21:18:30 GMT +0000 badg.us (badg.us) is DOWN since 2015-01-01 02:16:42 GMT +0000 badg.us (badg.us) is UP since 2015-01-01 02:18:55 GMT +0000 Should track down root cause and document.
Also had a cloudwatch alert saying that the host was down, so I'm thinking it was a network issue. Nothing in New Relic suggests errors with the Discourse app.
I too am deducing that AWS is at fault here - looks like a network blip. Would be a huge co-incidence if two sites went down in the same day but weren't related to the common platform, especially since we already know that one of the apps didn't fail. Should try and check badg.us logs though
Given that we're sunsetting badg.us, I think it's safe to close this. It's too late to diagnose why Discourse went down.
Status: NEW → RESOLVED
Closed: 11 years ago
Resolution: --- → WONTFIX

Bulk move of bugs

Component: Community IT: Infrastructure → MCWS
Product: Infrastructure & Operations → Participation Infrastructure
You need to log in before you can comment on or make changes to this bug.