Closed Bug 1510200 (t-yosemite-r7-226) Opened 7 years ago Closed 6 years ago

[MDC2] t-yosemite-r7-226 problem tracking

Categories

(Infrastructure & Operations :: RelOps: Hardware, defect)

defect
Not set
normal

Tracking

(Not tracked)

RESOLVED WORKSFORME

People

(Reporter: riman, Assigned: dividehex)

Details

This worker [1] fails almost any tasks. I have quarantined it for now. From the logs it looks like the worker has the same issues like t-yosemite-r7-386 Bug 1481050. [1] - https://tools.taskcluster.net/provisioners/releng-hardware/worker-types/gecko-t-osx-1010/workers/mdc2/t-yosemite-r7-226
Burns tests like a charm.. From 20 tests: - 1 Exception - 14 Fails - 5 Completed Does not look good. -> Quarantined. dhouse, please look into it when you have time Thank you!
Flags: needinfo?(dhouse)

I reimaged this machine today. It looks like it was last reimaged in July. So there may have been some broken or leftover files. If it keeps crashing, then we likely have a hardware failure.

It is under warranty until April.

I've removed the quarantine, and we can see if it crashes again.

Flags: needinfo?(dhouse)

re-quarantined. It failed two jobs.

[...]
Jan 09 02:44:58 t-yosemite-r7-226.test.releng.mdc2.mozilla.com sandboxd: ([2195]): plugin-container(2195) deny file-read-metadata /Users/cltbld/
Jan 09 02:45:26 t-yosemite-r7-226 kernel: process python[2181] caught causing excessive wakeups. Observed wakeups rate (per sec): 918; Maximum permitted wakeups rate (per sec): 150; Observation period: 300 seconds; Task lifetime number of wakeups: 45030
Jan 09 02:45:26 t-yosemite-r7-226 com.apple.xpc.launchd: (com.apple.ReportCrash[2201]): Endpoint has been activated through legacy launch(3) APIs. Please switch to XPC or bootstrap_check_in(): com.apple.ReportCrash
Jan 09 02:45:26 t-yosemite-r7-226 kernel: CODE SIGNING: cs_invalid_page(0x102de5000): p=2202[spindump] final status 0x2000000, allowing (remove VALID) page
Jan 09 02:45:26 t-yosemite-r7-226.test.releng.mdc2.mozilla.com ReportCrash: Invoking spindump for pid=2181 wakeups_rate=918 duration=50 because of excessive wakeups
Jan 09 02:45:26 t-yosemite-r7-226.test.releng.mdc2.mozilla.com SubmitDiagInfo: Couldn't load config file from on-disk location. Falling back to default location. Reason: Won't serialize in _readDictionaryFromJSONData due to nil object
Jan 09 02:45:26 t-yosemite-r7-226.test.releng.mdc2.mozilla.com spindump: Error grabbing microstackshots: 53
Jan 09 02:45:26 t-yosemite-r7-226.test.releng.mdc2.mozilla.com spindump: No microstackshots found
Jan 09 02:47:21 t-yosemite-r7-226.test.releng.mdc2.mozilla.com ReportCrash: Invoking spindump for pid=1649 wakeups_rate=267 duration=169 because of excessive wakeups
Jan 09 02:47:21 t-yosemite-r7-226.test.releng.mdc2.mozilla.com spindump: No microstackshots found
Jan 09 02:47:21 t-yosemite-r7-226.test.releng.mdc2.mozilla.com spindump: Error grabbing microstackshots: 53
Jan 09 02:48:55 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/09 02:48:55 About to reclaim task Lvr5V6xXS9CAxQgFz0AiIw...
Jan 09 02:48:55 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/09 02:48:55 Reclaiming task Lvr5V6xXS9CAxQgFz0AiIw... 
[...]
Jan 09 03:11:18 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/09 03:11:18 Resolving task Lvr5V6xXS9CAxQgFz0AiIw ...
Jan 09 03:11:18 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/09 03:11:18 ERROR(s) encountered: exit status 2
Jan 09 03:11:18 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/09 03:11:18 Removing task directory '/Users/cltbld/tasks/task_1547014345'... 

Maybe this machine is running slow. I didn't find full-machine crashes, but only the typical crashes within the tests. The tests are being marked as failed because of timing-out.

test-macosx64/debug-test-verify-e10s (https://tools.taskcluster.net/groups/TDgf4JylT_WCqsreChyMWg/tasks/Lvr5V6xXS9CAxQgFz0AiIw/runs/0)

03:01:27    ERROR - TEST-UNEXPECTED-TIMEOUT | automation.py | application timed out after 370 seconds with no output
03:01:27    ERROR - Force-terminating active process(es).
03:01:29  WARNING - runtests.py | Failed to get app exit code - running/crashed?
03:01:29    ERROR - TEST-UNEXPECTED-FAIL | ShutdownLeaks | process() called before end of test suite

test-macosx64/debug-web-platform-tests-e10s-9 (https://tools.taskcluster.net/groups/de5etKaQQu2XfmEWsznzjA/tasks/cpr5DSVHQHGh76vRt6AhoQ/runs/0)

Jan 08 22:09:12 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/08 22:09:12 Resolving task cpr5DSVHQHGh76vRt6AhoQ ...
Jan 08 22:09:13 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/08 22:09:13 ERROR(s) encountered: Task aborted - max run time exceeded 

CIDuty, could your team compare test timing to see if this machine is running slow? It looks like the tests are timing-out, and not crashing/failing like on other problem hardware.

I un-quarantined it again to let it run some other tests today for comparison.

Flags: needinfo?(ciduty)

Hello Dave,
I have looked into the worker and collected some timings for the tests it run. https://docs.google.com/spreadsheets/d/1rET1wTQSB0uR7nddxLhWsBr0-Eho2j6k8BX8pUjagP4/edit?usp=sharing

I also monitored the state of the while doing tasks H_Xya0hbSWCRDu_MXfvzlA task ID which I monitored and fcc_L36CSlGV6niZF_hLrw what I am monitoring right now.

I don't see anything weird looking into processes.

Flags: needinfo?(ciduty)

I had QTS reset the NVRAM/PRAM and SMC for this machine, since then it has run 7 tasks without problems and one that failed but failed for non-timeout reasons (appear to be a valid failure).

==> process 2031 will purposefully crash
TEST-UNEXPECTED-FAIL | leakcheck | tab missing output line for total leaks!

I'll leave it un-quarantined to see if it continues today without the many timeout failures we saw before.

It has successfully run 5 more tests and failed 1 (https://tools.taskcluster.net/groups/M3P_eETQRcGOotWPxjhXDA/tasks/UR6Rd103SNWUDCxf7_YCYA/runs/0) which looks like a valid failure (it is not a timeout). So, it appears to be working correctly now and not running as slow as yesterday/before.

4 more tasks are good

CIDuty, could you compare timing again on tests to see if this machine is working okay again?

Flags: needinfo?(ciduty)

Worker has been quarantined as it is completing jobs twice as slow compared to other OSX workers.

Flags: needinfo?(ciduty)

I've re-imaged the worked and removed it from quarantine.
We will keep on monitor it to see if the issue still happens.

QTS ran internet diagnostics and reset the SMC and NVRAM. No problems found
But it is still slow (and failing tests because of timeout).
I've quarantined it again and will get it un-racked for taking to an apple store for testing/repairs.

Dave, are there any news about the machine from the apple store?

Flags: needinfo?(dhouse)

I've ran a reimage over the machine and removed it from quarantine.
We will check it again, later on, when it will recover, if it takes tasks as intended without issues.

The machine is still faulty, back to quarantine!

Nothing was found to be wrong with this machine when it was taken to the apple store.

I've powered it off and will turn it back on in a few days to see if it come up after that and does not have problems/crashing.

Type: task → defect

removed machine from quarantine, rebooted and re-imaged.

currently, it looks like it didn't received multiple jobs in the same time. We will keep on monitoring it..

Flags: needinfo?(dhouse)

It still has issues: https://tools.taskcluster.net/provisioners/releng-hardware/worker-types/gecko-t-osx-1010/workers/mdc2/t-yosemite-r7-226

from the last failed tasks: [taskcluster:error] Task aborted - max run time exceeded

I've added it back to quarantine.

Hey Dave, i think this worker still needs a good checkup. What do you say?

Flags: needinfo?(dhouse)
Component: CIDuty → RelOps: Hardware
QA Contact: dlabici
Assignee: nobody → jwatkins

this worker appears to be working fine under mojave

Status: NEW → RESOLVED
Closed: 6 years ago
Flags: needinfo?(dhouse)
Resolution: --- → WORKSFORME
You need to log in before you can comment on or make changes to this bug.