Closed Bug 680494 Opened 15 years ago Closed 15 years ago

Network hiccups affected w32 slaves

Categories

(Release Engineering :: General, defect)

x86
Windows Server 2003
defect
Not set
normal

Tracking

(Not tracked)

RESOLVED FIXED

People

(Reporter: armenzg, Assigned: armenzg)

Details

Attachments

(1 file)

Attached file list of win32 slaves —
We had a lot and a lot of Win32 slaves that were logged in but due to network hiccups lost their ability to talk with buildbot and take jobs. I have attached the list and here is the list of slaves that got rebooted. mw32-ix-slave{16,17,22,23,24} w32-ix-slave{02,03,06,07,08,10,11,12,13,14,15,16,17,21,23} mw32-ix-slave{02,06,07,08,10,11,12,14} w32-ix-slave{25,26,29,30,33,34,35,37,40,41} I left w32-ix-slave05 undone in case anyone wants to investigate this and how to prevent it from happening when we have network nightmares. This is similar to bug 680457 but not related to OPSI from the few slaves a looked at.
Going through http://build.mozilla.org/builds/last-job-per-slave.txt prod-win32: mw32-ix-slave02 - loaner (bug 677520) mw32-ix-slave04 - failed to reboot on 2011-08-19, manually rebooted w32-ix-slave26 - config issues, bug 677888 w32-ix-slave27 - config issues, bug 677888 w32-ix-slave34 - in staging now, should update production_config.py w32-ix-slave41 - waiting for setup, bug 672973 w32-ix-slave42 - failed to reboot on 2011-08-09, manually rebooted try-win32: mw32-ix-slave22 - failed to reboot on 2011-08-15, manually rebooted mw32-ix-slave23 - failed to reboot on 2011-08-16, manually rebooted and kinda ran out of steam here. Probably need to look at try again once load goes up again at the start of the week.
The following list of slaves does not have to be due to the network hiccups but they need attention (perhaps I should file a different bug). This list has been obtained through this: grep w32 last-job-per-slave.txt | grep day | sort You can see some of the notes from slavealloc in here: mw32-ix-slave02 - loaned to jlebar; bug 677520 mw32-ix-slave22 - rebooted mw32-ix-slave23 - rebooted mw32-ix-slave24 - rebooted w32-ix-slave06 - dead - bug 673972 w32-ix-slave10 - rebooted w32-ix-slave14 - rebooted w32-ix-slave18 - rebooted w32-ix-slave20 - rebooted w32-ix-slave21 - rebooted w32-ix-slave26 - rebooted - slow (bug 663025) ? bug 677888 w32-ix-slave27 - rebooted - bug 677888 w32-ix-slave29 - rebooted - bug 677888 w32-ix-slave34 - rebooted w32-ix-slave41 - rebooted - bug 673436 After rebooting and checking twistd.log: All of them reconnected except #34 & $41 which I believe need to be setup from bug 673436. I should update the spreadsheet and make it defunct. I will check all of these slaves tomorrow and see if they took any jobs and it went.
Assignee: nobody → armenzg
Status: NEW → ASSIGNED
I didn't have time to look at this today but it looks like we are doing better. We are down to this list of slaves: mw32-ix-slave02 16 days, 1:37:26 w32-ix-slave03 1 day, 9:49:09 w32-ix-slave06 77 days, 1:05:55 w32-ix-slave34 328 days, 7:26:41 w32-ix-slave41 244 days, 6:11:17 I will deal with #34 and #41
I hit submit too fast: I didn't have time to look at this today but it looks like we are doing better. We are down to this list of slaves: mw32-ix-slave02 16 days, 1:37:26 w32-ix-slave03 1 day, 9:49:09 w32-ix-slave06 77 days, 1:05:55 w32-ix-slave34 328 days, 7:26:41 w32-ix-slave41 244 days, 6:11:17 I will deal with #34 and #41 in bug 682083. We should have better wait times tomorrow. Bug 681111 got fixed and should help detect slaves that have not done jobs recently.
Status: ASSIGNED → RESOLVED
Closed: 15 years ago
Resolution: --- → FIXED
Product: mozilla.org → Release Engineering
You need to log in before you can comment on or make changes to this bug.

Attachment

General

Created:
Updated:
Size: