Closed Bug 765493 Opened 14 years ago Closed 14 years ago

esx7.private.scl3 crashed, rebooted

Categories

(Infrastructure & Operations :: Virtualization, task)

task
Not set
normal

Tracking

(Not tracked)

RESOLVED FIXED

People

(Reporter: mburns, Assigned: gcox)

Details

Attachments

(1 file)

[06:40:37] <nagios-scl3> [544] vc1.private.scl3.mozilla.com:vmware_vcenter is CRITICAL: [esx7.private.scl3.mozilla.com] Host connection and power state is red BR [esx7.private.scl3.mozilla.com] vSphere HA host status is red esx7 crashed around 6:40am PST. Lerxst gave the all-clear to reboot it at 9:00am. Appears to be the same 'motherboard interupt' problem of previous hard-lockups.
Greg, please open a SR with VMware for this. We've seen this error many times before but never on the same host twice.
Assignee: server-ops → gcox
SR 12187737806 opened, target response is mid/late on the 19th.
Status: NEW → ASSIGNED
VMware got back with me. LINT1 is an NMI from hardware (which we knew). Per (many directions) we were suggested to upgrade firmware, which I've done in the SCL3 batch (PHX1 and HCI are untouched right now allowing for settle-out time). Upgraded stuff included onboard and HP nics, iLO, and the SmartArray. According to VMware, HP wants people to get up with them on LINT1's(!?). Working on getting an account set up to case this with them.
Upgraded PHX1's firmwares. HCI untouched. Opened HP case 4640647448; gathered requested logs.
The motherboard got replaced on 28 June. 1.5 passes of memtest came back clean. Going to do more burning over the weekend.
After the mobo swap, ran the box on a CPU burn for just under 3 days, no issues. Booted to ESX, ran at full CPU and 75% memory for a day, no issues. Having not seen anything for that long, calling it. Returned to service.
Status: ASSIGNED → RESOLVED
Closed: 14 years ago
Resolution: --- → FIXED
Product: mozilla.org → Infrastructure & Operations
You need to log in before you can comment on or make changes to this bug.

Attachment

General

Created:
Updated:
Size: