[Tracking] Treeherder jobs appear as running but they are completed.
Categories
(Infrastructure & Operations Graveyard :: CIDuty, task)
Tracking
(Not tracked)
People
(Reporter: bcrisan, Unassigned)
Details
This is a tracking bug for issue described in bug 1539009.
As part of the investigation:
It has been escalated to #ci
•bcrisan|ciduty> Hello, currently we have issues with jobs
01:37 https://treeherder.mozilla.org/#/jobs?repo=autoland&resultStatus=success%2Crunning%2Ctestfailed%2Cbusted%2Cexception%2Crunnable&classifiedState=unclassified&tochange=044a64c70a3b7072f1d9e00097b7a8745f43e709&fromchange=dde435d430ff4b6407833ea5f902febdc6a9497d&selectedJob=235933655&searchStr=linux%2Cx64%2Cpgo%2Cmochitests%2Cwith%2Ce10s%2Ctest-linux64-pgo%2Fopt
01:37 -mochitest-browser-chrome-e10s-4%2Cm-e10s%28bc4%29
01:37 is running for 286 minutes and it should run for just 25 minutes (at max)
01:37 same goes for many other tests and builds
01:38
<tomprince> bcrisan|ciduty: If you look at taskcluster (rather than treeherder) are they still running?
01:38
<•bcrisan|ciduty> checking
01:44 tomprince: looking into taskcluster the job looks completed, I have re-triggered it and it run for 20 seconds and marked it as completed
01:44 https://treeherder.mozilla.org/#/jobs?repo=autoland&resultStatus=success%2Crunning%2Ctestfailed%2Cbusted%2Cexception%2Crunnable&classifiedState=unclassified&tochange=044a64c70a3b7072f1d9e00097b7a8745f43e709&fromchange=dde435d430ff4b6407833ea5f902febdc6a9497d&selectedJob=235933655&searchStr=linux%2Cx64%2Cpgo%2Cmochitests%2Cwith%2Ce10s%2Ctest-linux64-pgo%2Fopt
01:44 https://tools.taskcluster.net/groups/UfIK1tVvR8W3-Uh3Ii9MCg/tasks/VJ7tKGw5RJSItOgfYQQwEw/runs/0/logs/public%2Flogs%2Flive_backing.log
01:44 check the live_backing.log
01:45 ⇐ Matti quit (Matti@moz-lv58ne.hsi16.unitymediagroup.de) Ping timeout: 121 seconds
01:45
<tomprince> bcrisan|ciduty: There was a pulse outage earlier today, and that would have caused notifications from taskcluster to treeherder to be dropped.
01:46
<•bcrisan|ciduty> I'm aware of that, is there something we could do to make them work as expected?
01:54
<tomprince> I don't think so.
and in #treeherder:
bcrisan|ciduty> Hi, after the pulse outage that happened earlier today, treeherder doesn't show the correct status for the jobs, as they appear to be completed in taskcluster but in treeherder the jobs appear to run for 300 minutes
02:47
<KWierso> bcrisan|ciduty: got a link to a push with jobs like that? my recent try push seems to be completing on time
02:48
<bcrisan|ciduty> did a bug with links in it
02:48 https://bugzilla.mozilla.org/show_bug.cgi?id=1539009
KWierso> bcrisan|ciduty: I think affected jobs are just never going to successfully resolve, since the pulse messages were lost
03:07
<bcrisan|ciduty> are they going to get some sort of time out at some point?
03:07 → aerickson joined (aerickson@moz-lsp.bt8.138.155.IP)
03:07
<KWierso> I can't remember how that works
03:08
<bcrisan|ciduty> no worries
The jobs triggered post event, does report correctly to treeherder and only those who got involved in the pulse outage are affected.
The trees were not closed, the issue it's not a blocker and everything (except those results) it's looking good.
Updated•7 years ago
|
Comment 1•7 years ago
|
||
Today we encountered a similar issue with this build: https://treeherder.mozilla.org/#/jobs?repo=mozilla-central&resultStatus=pending%2Crunning%2Ctestfailed%2Cbusted%2Cexception&classifiedState=unclassified&revision=7e40e33da3da2640e965a153254594a234231f76&selectedJob=242737451
Dustin confirmed that the pulse was working as expected, in that case there could be some issues from treeherder.
messages from IRC:
dustin> Dustin J Mitchell riman|ciduty: what are you seeing?
2:57 PM at a glance it looks fine to me
<•riman|ciduty> this build appear to be in pending for ~800 min but in TC seems to be completed: https://treeherder.mozilla.org/#/jobs?repo=mozilla-central&resultStatus=pending%2Crunning%2Ctestfailed%2Cbusted%2Cexception&classifiedState=unclassified&revision=7e40e33da3da2640e965a153254594a234231f76&selectedJob=242726618
<dustin> Dustin J Mitchell maybe a treeherder issue?
3:38 PM pulse has about 300 queued messages which is a very normal amount
3:38 PM looks like those are all for releng-services, too
<apavel|sheriffduty> Aryx: ^ do you know of any TH issues last night?
<dustin> Dustin J Mitchell treeherder-prod/jobs is getting 5-10 m/s and seems to be handling them
3:39 PM that's about as much as I can see
3:39 PM it's acking them anyway :)
<Aryx> https://tools.taskcluster.net/groups/JvRFBH7dSPyGZWXDAWEI7g/tasks/X1MdFRB7Q7GlyHpveeZpZw/runs/0 shows it completed
3:40 PM maybe some pulse messages got lost?
<dustin> Dustin J Mitchell hm, looks like there might have been 2 stuck messages until ~1h ago: https://irccloud.mozilla.com/file/gSWnUnu2/image.pngimage.png9.31KB • image/png
3:41 PM I'd guess something in the th consumption pipeline is failing?
3:43 PM oh, I think that's an artifact of the graph (I think it extrapolates missing data and is missing data for >1h ago)
3:43 PM tc-treeherder is also chugging along at about 10m/s
3:43 PM this all looks ok to me
| Reporter | ||
Comment 2•7 years ago
|
||
We should probably get this closed.
This is just a tracking bug for 1539009 that it's basically a duplicate of 1296077. Apparently they found out someone to look into the issue and get it fixed. (someone was assigned on 1296077 12 days ago).
| Reporter | ||
Updated•7 years ago
|
Updated•6 years ago
|
Description
•