Perma android fenix raptor-perftest Critical: Connection to Raptor webextension failed!
Categories
(Testing :: Raptor, defect, P3)
Tracking
(Not tracked)
People
(Reporter: intermittent-bug-filer, Unassigned)
References
Details
(Keywords: intermittent-failure, Whiteboard: [stockwell fixed:other])
User Story
Filed by: rmaries [at] mozilla.com
Parsed log: https://treeherder.mozilla.org/logviewer.html#?job_id=305579728&repo=autoland
Full log: https://firefox-ci-tc.services.mozilla.com/api/queue/v1/task/daSQlqnxRvmpgFy3HmoqnA/runs/0/artifacts/public/logs/live_backing.log
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - adb shell_output: adb -s HT7BN1A01909 wait-for-device shell su -c "chmod -R 777 /data/local/tmp/tests/raptor/profile/minidumps", timeout: None, root: True, timedout: None, exitcode: 0, output:
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - adb command_output: adb -s HT7BN1A01909 wait-for-device pull /data/local/tmp/tests/raptor/profile/minidumps /tmp/tmpM8FKVo/minidumps, timeout: None, timedout: None, exitcode: 0, output: /data/local/tmp/tests/raptor/profile/minidumps/: 0 files pulled, 0 skipped.
[task 2020-06-09T05:17:56.254Z] 05:17:46 CRITICAL - raptor-perftest Critical: Connection to Raptor webextension failed!
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-webext-android Info: removing reverse socket connections
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - adb command_output: adb -s HT7BN1A01909 wait-for-device reverse --remove-all, timeout: None, timedout: None, exitcode: 0, output:
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-webext-android Info: skipping check_for_crashes: application has not been launched
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-mitmproxy Info: Stopping mitmproxy playback, killing process 824
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-mitmproxy Info: Successfully killed the mitmproxy playback process
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-mitmproxy Info: Netlocs file is not available! Cant find /builds/task_1591679626/workspace/build/blobber_upload_dir/mitm_netlocs_mitm4-pixel2-fennec-amazon.json
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-webext Info: Mozproxy replay confidence data not available!
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-webext Info: removing webext /builds/task_1591679626/workspace/build/tests/raptor/raptor/webextension/../../webext/raptor
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-webext-android Info: removing test folder for raptor: /data/local/tmp/tests/raptor
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - adb shell_output: adb -s HT7BN1A01909 wait-for-device shell sync, timeout: None, root: False, timedout: None, exitcode: 0, output:
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - adb shell_output: adb -s HT7BN1A01909 wait-for-device shell su -c "rm -r /data/local/tmp/tests/raptor", timeout: None, root: True, timedout: None, exitcode: 0, output:
[task 2020-06-09T05:17:56.254Z] 05:17:47 INFO - adb shell_output: adb -s HT7BN1A01909 wait-for-device shell sync, timeout: None, root: False, timedout: None, exitcode: 0, output:
[task 2020-06-09T05:17:56.254Z] 05:17:47 INFO - adb shell_bool: adb -s HT7BN1A01909 wait-for-device shell su -c "test -e /data/local/tmp/tests/raptor", timeout: None, root: True, timedout: None, exitcode: 1, output:
[task 2020-06-09T05:17:56.254Z] 05:17:47 INFO - raptor-perftest Info: Removing temporary directory: /tmp/tmp1gde69
[task 2020-06-09T05:17:56.254Z] 05:17:47 INFO - raptor-perftest Info: Removing temporary directory: /tmp/tmpa6rYle
[task 2020-06-09T05:17:56.254Z] 05:17:47 INFO - raptor-control-server Info: shutting down control server
[task 2020-06-09T05:17:56.254Z] 05:17:47 INFO - raptor-webext Info: finished
[task 2020-06-09T05:17:56.254Z] 05:17:47 ERROR - Return code: 1
Comment 1•6 years ago
|
||
| Comment hidden (Intermittent Failures Robot) |
Comment 3•6 years ago
|
||
Likely an issue with the machine pool because retriggers of successful tasks fail: https://treeherder.mozilla.org/#/jobs?repo=mozilla-central&searchStr=android-hw-p2-8-0-android-aarch64-shippable%2Copt%2Craptor%2Cperformance%2Ctests%2Con%2Cfirefox%2Ctest-android-hw-p2-8-0-android-aarch64-shippable%2Fopt-raptor-tp6m-1-geckoview-cold-e10s%2Crap%28tp6m-c-1%29&revision=63dc5e9b1b02b0aebd6badfe5eaef7bb9aa8f430
Half the machine pool are machines which stopped working 16 hours ago. Could this be related? https://firefox-ci-tc.services.mozilla.com/provisioners/proj-autophone/worker-types/gecko-t-bitbar-gw-perf-p2
Updated•6 years ago
|
Comment 4•6 years ago
|
||
I don't think the idle devices is related to the raptor failure which looks like a "normal" type of raptor failure not related to the device though it could be related to a host connection issue. The idle device issue is a bitbar issue however and I expect that there is an issue with the hosts. aerickson is the go to person for bitbar.
Comment 5•6 years ago
|
||
Thank you. Greg, can you check what's going on with the raptor failures?
Comment 6•6 years ago
|
||
Half the machine pool are machines which stopped working 16 hours ago. Could this be related? https://firefox-ci-tc.services.mozilla.com/provisioners/proj-autophone/worker-types/gecko-t-bitbar-gw-perf-p2
They were moved to the unit queue. They appear in the old queue for 24 hours until they drop off. Everything seems ok in Bitbar land.
| Comment hidden (Intermittent Failures Robot) |
Comment 8•6 years ago
•
|
||
:aryx, the patch to disable all raptor geckoview pageload tests is landing today so this will drastically reduce the number of failures here.
That said, it's still an issue for the benchmark/resource-usage tests and it looks like it's possible that there's a commit that caused this regression - did you manage to find a commit/ci-change that may have caused this in your backfills?
Based on your comment about retriggers failing it sounds like this is a hardware issue and not something to do with any m-c changes. :aerickson, is it possible that something changed in bitbar or somewhere in CI on June 8th?
Updated•6 years ago
|
Comment 9•6 years ago
|
||
https://changelog.dev.mozaws.net/ lists all changes. On June 8, I tested a new Bitbar Docker image (but not in the production queues that these jobs would use). I don't think any other test infrastructure changes would affect test timing or success.
There doesn't seem to be one particularly bad host among the workers in that pool:
gecko-t-bitbar-gw-perf-p2.pixel2-57 {sr: [==== ] 44.4%, suc: 8, cmp: 18, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-54 {sr: [===== ] 50.0%, suc: 9, cmp: 18, exc: 1, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-55 {sr: [===== ] 50.0%, suc: 10, cmp: 20, exc: 0, rng: 0, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-56 {sr: [===== ] 52.6%, suc: 10, cmp: 19, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-59 {sr: [===== ] 52.6%, suc: 10, cmp: 19, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-53 {sr: [===== ] 57.9%, suc: 11, cmp: 19, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-58 {sr: [====== ] 61.1%, suc: 11, cmp: 18, exc: 1, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-49 {sr: [====== ] 63.2%, suc: 12, cmp: 19, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-50 {sr: [====== ] 63.2%, suc: 12, cmp: 19, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-51 {sr: [====== ] 68.4%, suc: 13, cmp: 19, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}
Comment 10•6 years ago
|
||
Thanks for checking :aerickson!
That's very odd how it just started all of a sudden. I wonder if it's the conditioned profiles causing this.
| Comment hidden (Intermittent Failures Robot) |
Comment 12•6 years ago
|
||
:tarek, would you be able to look into this issue? It seems like the conditioned profile for P2 aarch64 is busted when I run locally.
When I run without it, everything works fine, but as soon as the conditioned profile is used then the raptor webextension doesn't get installed and we fail with the "connection failed" issue.
Comment 13•6 years ago
|
||
This bug should be fixed now but let's leave it open for some time just in case.
Comment 14•6 years ago
|
||
No new failures sin 12th of June.
| Comment hidden (Intermittent Failures Robot) |
Comment 16•6 years ago
|
||
Oh we have to disable conditioned profiles in the fenix branch too.
| Comment hidden (Intermittent Failures Robot) |
Comment 18•6 years ago
|
||
(In reply to Greg Mierzwinski [:sparky] from comment #16)
Oh we have to disable conditioned profiles in the fenix branch too.
Do we have a bug to disable cond profiles?
Comment 19•6 years ago
|
||
No we don't, you can make a PR over there and then paste the link for it here.
Comment 20•6 years ago
|
||
| Comment hidden (Intermittent Failures Robot) |
| Comment hidden (Intermittent Failures Robot) |
Comment 23•6 years ago
|
||
(In reply to Intermittent Failures Robot from comment #22)
Platform breakdown:
- windows10-64-ref-hw-2017: 1
misclassified.
| Comment hidden (Intermittent Failures Robot) |
Comment 25•6 years ago
|
||
Shall the failures on the G5 get their own bug or the summary of this bug be updated?
Comment 26•6 years ago
|
||
Well, this bug has become a catch-all for the generic Connection to Raptor webextension failed errors. Might was well generalize the bug to include g5s as well. I'm not sure how to properly change the summary in order to keep the treeherder suggestions working for sheriffs though.
The current set of failures appear to be all fenix and power related. The test runs and then fails after attempting to pull the minidumps. Has the test actually completed and we are just cleaning up when the failure occurs or are we expecting to start another iteration?
Comment 27•6 years ago
|
||
There are no other tests supposed to run after the failure.
From a successful log:
INFO - adb command_output: adb -s ZY3223BKR3 wait-for-device bugreport /builds/task_1596199237/workspace/build/blobber_upload_dir, timeout: None, timedout: None, exitcode: 0, output: Bugreport is in progress and it could take minutes to complete.
INFO - Please be patient and do not cancel or disconnect your device until it completes.
INFO - /data/user_de/0/com.android.shell/files/bugreports/bugreport-NPP25.137-15-2020-07-31-12-50-37.zip: 1 file pulled, 0 skipped. 10.4 MB/s (1922967 bytes in 0.176s)
INFO - raptor-webext-android Info: removing reverse socket connections
INFO - adb command_output: adb -s ZY3223BKR3 wait-for-device reverse --remove-all, timeout: None, timedout: None, exitcode: 0, output:
The last two lines are missing from the log with the failure.
Greg, is this already covered by a different bug? Last successful run was on July 31st, first failed on August 3rd.
Comment 28•6 years ago
|
||
It looks like I can reproduce locally even without fission. I'll try to bisect and see if something jumps out.
Comment 29•6 years ago
|
||
I've raised this issue with the Fenix team as well since it's a fenix-only permafail: https://github.com/mozilla-mobile/fenix/issues/13399
Comment 30•6 years ago
|
||
So it's covered by a github issue, but not by another bug.
Comment 31•6 years ago
|
||
Are the timeouts seen here ([taskcluster:error] Task aborted - max run time exceeded) related to the issues discussed in this bug?
https://treeherder.mozilla.org/#/jobs?repo=mozilla-central&group_state=expanded&searchStr=android%2C7.0%2Cmotog5%2Cshippable%2Copt%2Craptor%2Cperformance%2Ctests%2Con%2Cfenix%2Ctest-android-hw-g5-7-0-arm7-api-16-shippable%2Fopt-raptor-scn-power-idle-fenix-e10s%2Cidl-p&fromchange=b579c12907dc3bb210177454a696ccc181c1ded0&selectedTaskRun=EzqG9ll7T8qF9KtCX3y21A.0&tochange=605c404fbd80c67e1127ac054b8bec6742bd748f
Or should we make a new bug to track those?
| Comment hidden (Intermittent Failures Robot) |
| Comment hidden (Intermittent Failures Robot) |
Comment 34•6 years ago
|
||
The resolution in the github issue is to port these tests asap so I'm going to look into porting these tests to raptor-browsertime this week. (leaving ni open for myself)
| Comment hidden (Intermittent Failures Robot) |
| Comment hidden (Intermittent Failures Robot) |
Comment 37•6 years ago
|
||
This has been fixed now - the power tests were migrated to browsertime.
Description
•