[MDC2] t-yosemite-r7-226 problem tracking
Categories
(Infrastructure & Operations :: RelOps: Hardware, defect)
Tracking
(Not tracked)
People
(Reporter: riman, Assigned: dividehex)
Details
Comment 1•7 years ago
|
||
I reimaged this machine today. It looks like it was last reimaged in July. So there may have been some broken or leftover files. If it keeps crashing, then we likely have a hardware failure.
It is under warranty until April.
I've removed the quarantine, and we can see if it crashes again.
[...]
Jan 09 02:44:58 t-yosemite-r7-226.test.releng.mdc2.mozilla.com sandboxd: ([2195]): plugin-container(2195) deny file-read-metadata /Users/cltbld/
Jan 09 02:45:26 t-yosemite-r7-226 kernel: process python[2181] caught causing excessive wakeups. Observed wakeups rate (per sec): 918; Maximum permitted wakeups rate (per sec): 150; Observation period: 300 seconds; Task lifetime number of wakeups: 45030
Jan 09 02:45:26 t-yosemite-r7-226 com.apple.xpc.launchd: (com.apple.ReportCrash[2201]): Endpoint has been activated through legacy launch(3) APIs. Please switch to XPC or bootstrap_check_in(): com.apple.ReportCrash
Jan 09 02:45:26 t-yosemite-r7-226 kernel: CODE SIGNING: cs_invalid_page(0x102de5000): p=2202[spindump] final status 0x2000000, allowing (remove VALID) page
Jan 09 02:45:26 t-yosemite-r7-226.test.releng.mdc2.mozilla.com ReportCrash: Invoking spindump for pid=2181 wakeups_rate=918 duration=50 because of excessive wakeups
Jan 09 02:45:26 t-yosemite-r7-226.test.releng.mdc2.mozilla.com SubmitDiagInfo: Couldn't load config file from on-disk location. Falling back to default location. Reason: Won't serialize in _readDictionaryFromJSONData due to nil object
Jan 09 02:45:26 t-yosemite-r7-226.test.releng.mdc2.mozilla.com spindump: Error grabbing microstackshots: 53
Jan 09 02:45:26 t-yosemite-r7-226.test.releng.mdc2.mozilla.com spindump: No microstackshots found
Jan 09 02:47:21 t-yosemite-r7-226.test.releng.mdc2.mozilla.com ReportCrash: Invoking spindump for pid=1649 wakeups_rate=267 duration=169 because of excessive wakeups
Jan 09 02:47:21 t-yosemite-r7-226.test.releng.mdc2.mozilla.com spindump: No microstackshots found
Jan 09 02:47:21 t-yosemite-r7-226.test.releng.mdc2.mozilla.com spindump: Error grabbing microstackshots: 53
Jan 09 02:48:55 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/09 02:48:55 About to reclaim task Lvr5V6xXS9CAxQgFz0AiIw...
Jan 09 02:48:55 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/09 02:48:55 Reclaiming task Lvr5V6xXS9CAxQgFz0AiIw...
[...]
Jan 09 03:11:18 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/09 03:11:18 Resolving task Lvr5V6xXS9CAxQgFz0AiIw ...
Jan 09 03:11:18 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/09 03:11:18 ERROR(s) encountered: exit status 2
Jan 09 03:11:18 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/09 03:11:18 Removing task directory '/Users/cltbld/tasks/task_1547014345'...
Maybe this machine is running slow. I didn't find full-machine crashes, but only the typical crashes within the tests. The tests are being marked as failed because of timing-out.
test-macosx64/debug-test-verify-e10s (https://tools.taskcluster.net/groups/TDgf4JylT_WCqsreChyMWg/tasks/Lvr5V6xXS9CAxQgFz0AiIw/runs/0)
03:01:27 ERROR - TEST-UNEXPECTED-TIMEOUT | automation.py | application timed out after 370 seconds with no output
03:01:27 ERROR - Force-terminating active process(es).
03:01:29 WARNING - runtests.py | Failed to get app exit code - running/crashed?
03:01:29 ERROR - TEST-UNEXPECTED-FAIL | ShutdownLeaks | process() called before end of test suite
test-macosx64/debug-web-platform-tests-e10s-9 (https://tools.taskcluster.net/groups/de5etKaQQu2XfmEWsznzjA/tasks/cpr5DSVHQHGh76vRt6AhoQ/runs/0)
Jan 08 22:09:12 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/08 22:09:12 Resolving task cpr5DSVHQHGh76vRt6AhoQ ...
Jan 08 22:09:13 t-yosemite-r7-226.test.releng.mdc2.mozilla.com generic-worker: 2019/01/08 22:09:13 ERROR(s) encountered: Task aborted - max run time exceeded
CIDuty, could your team compare test timing to see if this machine is running slow? It looks like the tests are timing-out, and not crashing/failing like on other problem hardware.
I un-quarantined it again to let it run some other tests today for comparison.
Comment 7•7 years ago
|
||
Hello Dave,
I have looked into the worker and collected some timings for the tests it run. https://docs.google.com/spreadsheets/d/1rET1wTQSB0uR7nddxLhWsBr0-Eho2j6k8BX8pUjagP4/edit?usp=sharing
I also monitored the state of the while doing tasks H_Xya0hbSWCRDu_MXfvzlA task ID which I monitored and fcc_L36CSlGV6niZF_hLrw what I am monitoring right now.
I don't see anything weird looking into processes.
I had QTS reset the NVRAM/PRAM and SMC for this machine, since then it has run 7 tasks without problems and one that failed but failed for non-timeout reasons (appear to be a valid failure).
==> process 2031 will purposefully crash
TEST-UNEXPECTED-FAIL | leakcheck | tab missing output line for total leaks!
I'll leave it un-quarantined to see if it continues today without the many timeout failures we saw before.
It has successfully run 5 more tests and failed 1 (https://tools.taskcluster.net/groups/M3P_eETQRcGOotWPxjhXDA/tasks/UR6Rd103SNWUDCxf7_YCYA/runs/0) which looks like a valid failure (it is not a timeout). So, it appears to be working correctly now and not running as slow as yesterday/before.
Comment 10•7 years ago
|
||
4 more tasks are good
CIDuty, could you compare timing again on tests to see if this machine is working okay again?
Comment 11•7 years ago
|
||
Worker has been quarantined as it is completing jobs twice as slow compared to other OSX workers.
Comment 12•7 years ago
|
||
I've re-imaged the worked and removed it from quarantine.
We will keep on monitor it to see if the issue still happens.
Comment 13•7 years ago
|
||
QTS ran internet diagnostics and reset the SMC and NVRAM. No problems found
But it is still slow (and failing tests because of timeout).
I've quarantined it again and will get it un-racked for taking to an apple store for testing/repairs.
Comment 14•7 years ago
|
||
Dave, are there any news about the machine from the apple store?
Comment 15•7 years ago
|
||
I've ran a reimage over the machine and removed it from quarantine.
We will check it again, later on, when it will recover, if it takes tasks as intended without issues.
Comment 16•7 years ago
|
||
The machine is still faulty, back to quarantine!
Comment 17•7 years ago
|
||
Nothing was found to be wrong with this machine when it was taken to the apple store.
I've powered it off and will turn it back on in a few days to see if it come up after that and does not have problems/crashing.
Updated•7 years ago
|
Comment 18•7 years ago
|
||
removed machine from quarantine, rebooted and re-imaged.
Comment 19•7 years ago
|
||
currently, it looks like it didn't received multiple jobs in the same time. We will keep on monitoring it..
| Reporter | ||
Comment 20•7 years ago
|
||
It still has issues: https://tools.taskcluster.net/provisioners/releng-hardware/worker-types/gecko-t-osx-1010/workers/mdc2/t-yosemite-r7-226
from the last failed tasks: [taskcluster:error] Task aborted - max run time exceeded
I've added it back to quarantine.
Comment 21•7 years ago
|
||
Hey Dave, i think this worker still needs a good checkup. What do you say?
| Assignee | ||
Updated•7 years ago
|
| Assignee | ||
Updated•7 years ago
|
Comment 22•6 years ago
|
||
this worker appears to be working fine under mojave
Description
•