Android raptor jobs backlog and they end up as exception with deadline-exceeded
Categories
(Testing :: Raptor, defect, P1)
Tracking
(Not tracked)
People
(Reporter: CosminS, Assigned: aerickson)
Details
(Keywords: intermittent-failure)
Jobs are queued up for a long time (eg. Not started (queued for 1443 minute(s))
This happens so far only on beta:
https://treeherder.mozilla.org/#/jobs?repo=mozilla-beta&resultStatus=pending%2Crunning%2Cexception&tochange=ded211603671be4210c80f5c2f17c0a4a5f9ceb7&fromchange=a08cdd143e57dbd4bc062851b790ef53d82196ed&group_state=expanded&selectedJob=242335774
Taskcluster eg: https://tools.taskcluster.net/groups/IZcgnS5-TG61M9uxes40bQ/tasks/V8mkjo1_Rs-KafW3gI_uKA/runs/0
autoland, inbound and central are ok https://treeherder.mozilla.org/#/jobs?repo=mozilla-central&resultStatus=success%2Cpending%2Crunning%2Ctestfailed%2Cbusted%2Cexception&classifiedState=unclassified&searchStr=android%2Craptor&group_state=expanded&fromchange=8f9398c8c37d81bd0c3816ac852c475ac802cabf&tochange=ff3ec9547e4f15209af87010ed39a0590c253efe
From discussion on irc on #ci channel the worker pool are
gecko-t-ap-perf-p2
gecko-t-ap-perf-g5
https://tools.taskcluster.net/provisioners/proj-autophone/worker-types
| Assignee | ||
Comment 1•7 years ago
|
||
This is related to Bug 1474897.
I've moved more workers to the old tc-w perf queues. Please let me know if you have more failures.
| Assignee | ||
Updated•7 years ago
|
Updated•7 years ago
|
| Assignee | ||
Comment 2•7 years ago
|
||
update: The congestion on the old queues was due to Fenix jobs. I worked with mhentges to move the work to the new queues and things have been smooth since.
| Assignee | ||
Comment 3•7 years ago
|
||
Closing this as we seem to have addressed the initial issue.
The perf queues remain heavily saturated, but it's not for the original cause of this bug.
Description
•