Closed Bug 1644331 Opened 6 years ago Closed 6 years ago

Perma android fenix raptor-perftest Critical: Connection to Raptor webextension failed!

Categories

(Testing :: Raptor, defect, P3)

defect

Tracking

(Not tracked)

RESOLVED FIXED

People

(Reporter: intermittent-bug-filer, Unassigned)

References

Details

(Keywords: intermittent-failure, Whiteboard: [stockwell fixed:other])

User Story

https://github.com/mozilla-mobile/fenix/issues/13399

Filed by: rmaries [at] mozilla.com
Parsed log: https://treeherder.mozilla.org/logviewer.html#?job_id=305579728&repo=autoland
Full log: https://firefox-ci-tc.services.mozilla.com/api/queue/v1/task/daSQlqnxRvmpgFy3HmoqnA/runs/0/artifacts/public/logs/live_backing.log


[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - adb shell_output: adb -s HT7BN1A01909 wait-for-device shell su -c "chmod -R 777 /data/local/tmp/tests/raptor/profile/minidumps", timeout: None, root: True, timedout: None, exitcode: 0, output:
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - adb command_output: adb -s HT7BN1A01909 wait-for-device pull /data/local/tmp/tests/raptor/profile/minidumps /tmp/tmpM8FKVo/minidumps, timeout: None, timedout: None, exitcode: 0, output: /data/local/tmp/tests/raptor/profile/minidumps/: 0 files pulled, 0 skipped.
[task 2020-06-09T05:17:56.254Z] 05:17:46 CRITICAL - raptor-perftest Critical: Connection to Raptor webextension failed!
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-webext-android Info: removing reverse socket connections
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - adb command_output: adb -s HT7BN1A01909 wait-for-device reverse --remove-all, timeout: None, timedout: None, exitcode: 0, output:
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-webext-android Info: skipping check_for_crashes: application has not been launched
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-mitmproxy Info: Stopping mitmproxy playback, killing process 824
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-mitmproxy Info: Successfully killed the mitmproxy playback process
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-mitmproxy Info: Netlocs file is not available! Cant find /builds/task_1591679626/workspace/build/blobber_upload_dir/mitm_netlocs_mitm4-pixel2-fennec-amazon.json
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-webext Info: Mozproxy replay confidence data not available!
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-webext Info: removing webext /builds/task_1591679626/workspace/build/tests/raptor/raptor/webextension/../../webext/raptor
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - raptor-webext-android Info: removing test folder for raptor: /data/local/tmp/tests/raptor
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - adb shell_output: adb -s HT7BN1A01909 wait-for-device shell sync, timeout: None, root: False, timedout: None, exitcode: 0, output:
[task 2020-06-09T05:17:56.254Z] 05:17:46 INFO - adb shell_output: adb -s HT7BN1A01909 wait-for-device shell su -c "rm -r /data/local/tmp/tests/raptor", timeout: None, root: True, timedout: None, exitcode: 0, output:
[task 2020-06-09T05:17:56.254Z] 05:17:47 INFO - adb shell_output: adb -s HT7BN1A01909 wait-for-device shell sync, timeout: None, root: False, timedout: None, exitcode: 0, output:
[task 2020-06-09T05:17:56.254Z] 05:17:47 INFO - adb shell_bool: adb -s HT7BN1A01909 wait-for-device shell su -c "test -e /data/local/tmp/tests/raptor", timeout: None, root: True, timedout: None, exitcode: 1, output:
[task 2020-06-09T05:17:56.254Z] 05:17:47 INFO - raptor-perftest Info: Removing temporary directory: /tmp/tmp1gde69
[task 2020-06-09T05:17:56.254Z] 05:17:47 INFO - raptor-perftest Info: Removing temporary directory: /tmp/tmpa6rYle
[task 2020-06-09T05:17:56.254Z] 05:17:47 INFO - raptor-control-server Info: shutting down control server
[task 2020-06-09T05:17:56.254Z] 05:17:47 INFO - raptor-webext Info: finished
[task 2020-06-09T05:17:56.254Z] 05:17:47 ERROR - Return code: 1

Flags: needinfo?(jmaher)
Flags: needinfo?(gbrown)

I don't think the idle devices is related to the raptor failure which looks like a "normal" type of raptor failure not related to the device though it could be related to a host connection issue. The idle device issue is a bitbar issue however and I expect that there is an issue with the hosts. aerickson is the go to person for bitbar.

Flags: needinfo?(aerickson)

Thank you. Greg, can you check what's going on with the raptor failures?

Flags: needinfo?(jmaher)
Flags: needinfo?(gmierz2)
Flags: needinfo?(gbrown)

Half the machine pool are machines which stopped working 16 hours ago. Could this be related? https://firefox-ci-tc.services.mozilla.com/provisioners/proj-autophone/worker-types/gecko-t-bitbar-gw-perf-p2

They were moved to the unit queue. They appear in the old queue for 24 hours until they drop off. Everything seems ok in Bitbar land.

Flags: needinfo?(aerickson)
Depends on: 1644787

:aryx, the patch to disable all raptor geckoview pageload tests is landing today so this will drastically reduce the number of failures here.

That said, it's still an issue for the benchmark/resource-usage tests and it looks like it's possible that there's a commit that caused this regression - did you manage to find a commit/ci-change that may have caused this in your backfills?

Based on your comment about retriggers failing it sounds like this is a hardware issue and not something to do with any m-c changes. :aerickson, is it possible that something changed in bitbar or somewhere in CI on June 8th?

Flags: needinfo?(gmierz2) → needinfo?(aryx.bugmail)
Flags: needinfo?(aerickson)

https://changelog.dev.mozaws.net/ lists all changes. On June 8, I tested a new Bitbar Docker image (but not in the production queues that these jobs would use). I don't think any other test infrastructure changes would affect test timing or success.

There doesn't seem to be one particularly bad host among the workers in that pool:

gecko-t-bitbar-gw-perf-p2.pixel2-57 {sr: [==== ] 44.4%, suc: 8, cmp: 18, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-54 {sr: [===== ] 50.0%, suc: 9, cmp: 18, exc: 1, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-55 {sr: [===== ] 50.0%, suc: 10, cmp: 20, exc: 0, rng: 0, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-56 {sr: [===== ] 52.6%, suc: 10, cmp: 19, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-59 {sr: [===== ] 52.6%, suc: 10, cmp: 19, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-53 {sr: [===== ] 57.9%, suc: 11, cmp: 19, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-58 {sr: [====== ] 61.1%, suc: 11, cmp: 18, exc: 1, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-49 {sr: [====== ] 63.2%, suc: 12, cmp: 19, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-50 {sr: [====== ] 63.2%, suc: 12, cmp: 19, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}
gecko-t-bitbar-gw-perf-p2.pixel2-51 {sr: [====== ] 68.4%, suc: 13, cmp: 19, exc: 0, rng: 1, alerts: ['Low health (less than 0.85)!']}

Flags: needinfo?(aerickson)

Thanks for checking :aerickson!

That's very odd how it just started all of a sudden. I wonder if it's the conditioned profiles causing this.

:tarek, would you be able to look into this issue? It seems like the conditioned profile for P2 aarch64 is busted when I run locally.

When I run without it, everything works fine, but as soon as the conditioned profile is used then the raptor webextension doesn't get installed and we fail with the "connection failed" issue.

Flags: needinfo?(tarek)
Depends on: 1645135

This bug should be fixed now but let's leave it open for some time just in case.

Flags: needinfo?(tarek)
Flags: needinfo?(aryx.bugmail)
See Also: → 1645181

No new failures sin 12th of June.

Whiteboard: [stockwell disable-recommended] → [stockwell fixed:other]

Oh we have to disable conditioned profiles in the fenix branch too.

(In reply to Greg Mierzwinski [:sparky] from comment #16)

Oh we have to disable conditioned profiles in the fenix branch too.

Do we have a bug to disable cond profiles?

Flags: needinfo?(gmierz2)

No we don't, you can make a PR over there and then paste the link for it here.

Flags: needinfo?(gmierz2)

(In reply to Intermittent Failures Robot from comment #22)

Platform breakdown:

  • windows10-64-ref-hw-2017: 1

misclassified.

Shall the failures on the G5 get their own bug or the summary of this bug be updated?

Flags: needinfo?(bob)

Well, this bug has become a catch-all for the generic Connection to Raptor webextension failed errors. Might was well generalize the bug to include g5s as well. I'm not sure how to properly change the summary in order to keep the treeherder suggestions working for sheriffs though.

The current set of failures appear to be all fenix and power related. The test runs and then fails after attempting to pull the minidumps. Has the test actually completed and we are just cleaning up when the failure occurs or are we expecting to start another iteration?

Flags: needinfo?(bob)

There are no other tests supposed to run after the failure.

From a successful log:

 INFO -  adb command_output: adb -s ZY3223BKR3 wait-for-device bugreport /builds/task_1596199237/workspace/build/blobber_upload_dir, timeout: None, timedout: None, exitcode: 0, output: Bugreport is in progress and it could take minutes to complete.
 INFO -  Please be patient and do not cancel or disconnect your device until it completes.
 INFO -  /data/user_de/0/com.android.shell/files/bugreports/bugreport-NPP25.137-15-2020-07-31-12-50-37.zip: 1 file pulled, 0 skipped. 10.4 MB/s (1922967 bytes in 0.176s)
 INFO -  raptor-webext-android Info: removing reverse socket connections
 INFO -  adb command_output: adb -s ZY3223BKR3 wait-for-device reverse --remove-all, timeout: None, timedout: None, exitcode: 0, output:

The last two lines are missing from the log with the failure.

Greg, is this already covered by a different bug? Last successful run was on July 31st, first failed on August 3rd.

Flags: needinfo?(gmierz2)
Summary: Perma android-hw-p2-8-0 raptor-perftest Critical: Connection to Raptor webextension failed! → Perma android fenix raptor-perftest Critical: Connection to Raptor webextension failed!

It looks like I can reproduce locally even without fission. I'll try to bisect and see if something jumps out.

I've raised this issue with the Fenix team as well since it's a fenix-only permafail: https://github.com/mozilla-mobile/fenix/issues/13399

User Story: (updated)
Flags: needinfo?(gmierz2)

So it's covered by a github issue, but not by another bug.

The resolution in the github issue is to port these tests asap so I'm going to look into porting these tests to raptor-browsertime this week. (leaving ni open for myself)

Depends on: 1661329

This has been fixed now - the power tests were migrated to browsertime.

Status: NEW → RESOLVED
Closed: 6 years ago
Flags: needinfo?(gmierz2)
Resolution: --- → FIXED
You need to log in before you can comment on or make changes to this bug.