Closed Bug 1050510 Opened 12 years ago Closed 8 years ago

High load of CPU (95-99%) after run any app

Categories

(Firefox OS Graveyard :: Gaia::System, defect)

defect
Not set
normal

Tracking

(Not tracked)

RESOLVED WONTFIX

People

(Reporter: zrzut01, Unassigned)

References

Details

(Keywords: regression)

Attachments

(2 files)

User Agent: Mozilla/5.0 (X11; Linux x86_64; rv:30.0) Gecko/20100101 Firefox/30.0 (Beta/Release) Build ID: 20140605174243 Steps to reproduce: STR: 1. Reboot the device 2. Open an app, i.e. Settings or Dialer. Actual results: CPU load jump to 95-99% even after close the app. example from 'top -m 5': User 87%, System 12%, IOW 0%, IRQ 0% User 274 + Nice 0 + Sys 40 + Idle 0 + IOW 0 + IRQ 0 + SIRQ 0 = 314 PID PR CPU% S #THR VSS RSS PCY UID Name 138 0 97% R 48 178136K 80632K fg root /system/b2g/b2g 457 0 1% R 1 1024K 420K fg root top 111 0 0% S 5 4480K 188K fg root /sbin/adbd 4 0 0% S 1 0K 0K fg root kworker/0:0 5 0 0% D 1 0K 0K fg root kworker/u:0 The issue is reproducible almost every time (7 per 10 tests) but sometimes it occurs after a while. Expected results: CPU load can jump for a while but can't stay so high because currently it kills overall performance.
Found on: Alcatel One Touch Fire production (got from T-mobile Poland) B2G version: 2.1.0.0-prerelease master Platform version: 34.0a1 Build Identifier: 20140807084340 Git commit info: 2014-08-07 09:35:39 54c3c19d
Flags: needinfo?(dflanagan)
I'm seeing something similar since a couple of days on my Nexus S, with a constant >50% load when idling. PID 300 is Homescreen: User 60%, System 11%, IOW 0%, IRQ 0% User 10 + Nice 56 + Sys 12 + Idle 31 + IOW 0 + IRQ 0 + SIRQ 0 = 109 PID PR CPU% S #THR VSS RSS PCY UID Name 300 0 50% S 16 89540K 48244K fg app_300 /system/b2g/plugin-container 77 0 13% S 62 212488K 114360K fg root /system/b2g/b2g 588 0 3% R 1 996K 408K fg shell top 42 0 0% S 1 0K 0K fg root kworker/u:1 72 0 0% S 1 0K 0K fg root kworker/0:2 User 54%, System 10%, IOW 0%, IRQ 0% User 9 + Nice 49 + Sys 11 + Idle 38 + IOW 0 + IRQ 0 + SIRQ 0 = 107 PID PR CPU% S #THR VSS RSS PCY UID Name 300 0 49% S 16 89540K 48244K fg app_300 /system/b2g/plugin-container 77 0 12% S 62 212488K 114360K fg root /system/b2g/b2g 588 0 3% R 1 996K 408K fg shell top 72 0 0% S 1 0K 0K fg root kworker/0:2 5 0 0% S 1 0K 0K fg root kworker/u:0 User 51%, System 12%, IOW 0%, IRQ 0% User 10 + Nice 45 + Sys 13 + Idle 39 + IOW 0 + IRQ 0 + SIRQ 0 = 107 PID PR CPU% S #THR VSS RSS PCY UID Name 300 0 42% S 16 89540K 48244K fg app_300 /system/b2g/plugin-container 77 0 14% S 62 212488K 114360K fg root /system/b2g/b2g 588 0 4% R 1 996K 408K fg shell top 42 0 1% S 1 0K 0K fg root kworker/u:1
Could we get verification from QA on buri ?
I may be reproducing something similar on my Flame. For now, I have it more or less crashing when configured with 512M of RAM, only one CPU enabled, and not plugged on USB.
And switching to a kernel with more than 1 CPUs it's okay.
So far my device is now in good shape. It turned out to be bug 887198 which regressed.
(In reply to Alexandre LISSY :gerard-majax from comment #5) > Could we get verification from QA on buri ? To my understanding, Buri is not officially supported on 2.1 per product. We can check this on Flame though - let's branch check first here on Flame to determine if we can reproduce & find out if it's a regression.
Profile has been captured when there was ~72% continuous load of 138 PID process when nothing was running and device was idling. Profile link: http://people.mozilla.org/~bgirard/cleopatra/#report=58c77fd594330200f270842250d02cbc68f67717 Captured on: Alcatel One Touch Fire production (got from T-mobile Poland) B2G version: 2.1.0.0-prerelease master Platform version: 34.0a1 Build Identifier: 20140820162605 Git commit info: 2014-08-20 09:35:53 05768c07
This issue does NOT repro on Flame 2.1, Flame 2.0 After reboot and with Settings app open, running 'adb shell top -m 5' returns the following result: On 2.1 Flame: User 28%, System 26%, IOW 0%, IRQ 0% User 181 + Nice 15 + Sys 185 + Idle 305 + IOW 3 + IRQ 0 + SIRQ 4 = 693 PID PR CPU% S #THR VSS RSS PCY UID Name 295 1 30% S 60 212796K 91740K root /system/b2g/b2g 1586 1 3% S 18 83072K 34072K u0_a1586 /system/b2g/b2g 285 0 2% S 7 23804K 7508K fg media /system/bin/mediaserver 1333 1 2% S 19 70548K 27636K u0_a1333 /system/b2g/b2g 96 1 1% S 1 0K 0K root ksmd ---------- On 2.0 Flame: User 6%, System 8%, IOW 0%, IRQ 0% User 14 + Nice 4 + Sys 26 + Idle 253 + IOW 0 + IRQ 0 + SIRQ 0 = 297 PID PR CPU% S #THR VSS RSS PCY UID Name 1179 0 5% S 14 83472K 31484K u0_a1179 /system/b2g/plugin-container 96 0 4% S 1 0K 0K root ksmd 1420 0 3% R 1 1232K 560K root top 289 0 3% S 57 212944K 92876K root /system/b2g/b2g 359 0 0% S 10 11700K 896K radio /system/bin/qmuxd ---------- On 1.4 Flame: User 2%, System 4%, IOW 0%, IRQ 0% User 9 + Nice 0 + Sys 14 + Idle 283 + IOW 0 + IRQ 0 + SIRQ 0 = 306 PID PR CPU% S #THR VSS RSS PCY UID Name 1221 0 3% R 1 1232K 560K root top 293 0 1% S 43 207052K 83272K root /system/b2g/b2g 96 0 0% S 1 0K 0K root ksmd 400 0 0% S 5 4512K 228K root /sbin/adbd 982 0 0% S 5 6284K 624K root /system/bin/mpdecision ---------------------------- Tested on: Device: Flame BuildID: 20140828040749 Gaia: 39cad6c82122b964f12a66771bfbcc14fb342d9e Gecko: 2a15dc07ddaa Version: 34.0a1 (2.1 Master) Firmware: V123 User Agent: Mozilla/5.0 (Mobile; rv:33.0) Gecko/33.0 Firefox/33.0 Device: Flame BuildID: 20140828000650 Gaia: a6fc290a5601183f84ee9c7cb37eeebc933af2f5 Gecko: 625dd5529548 Version: 32.0 (2.0) Firmware: V123 User Agent: Mozilla/5.0 (Mobile; rv:32.0) Gecko/32.0 Firefox/32.0 Device: Flame BuildID: 20140827090228 Gaia: 05653cb12d324649687dad3eeb2ea373a2ad84d4 Gecko: baf01c5965ef Version: 30.0 (1.4) Firmware: V123 User Agent: Mozilla/5.0 (Mobile; rv:30.0) Gecko/30.0 Firefox/30.0
QA Whiteboard: [QAnalyst-Triage?]
Flags: needinfo?(jmitchell)
Keywords: qawanted
On first sentence of above comment I forgot to add this does NOT occur to Flame 1.4.
QA Whiteboard: [QAnalyst-Triage?] → [QAnalyst-Triage+]
Flags: needinfo?(jmitchell)
Issue still occurs on: Alcatel One Touch Fire production (got from T-mobile Poland) B2G version: 2.1.0.0-prerelease master Platform version: 34.0a1 Build Identifier: 20140828160238 Git commit info: 2014-08-28 15:45:33 007f3c50
After manual restart of gaia by: adb shell stop b2g && adb shell start b2g then continuous CPU load fall to 40%.
On my Alcatel One Touch Fire, i have about 40% CPU load all time B2G version: 2.1.0.0-prerelease master Platform version: 34.0a1 Build Identifier: 20140830234901 Git commit info: 2014-08-29 13:39:41 32b849d2
I have compiled my own new build, there is about 40% continuous CPU load. Alcatel One Touch Fire production (got from T-mobile Poland) B2G version: 2.1.0.0-prerelease master Platform version: 34.0a1 Build Identifier: 20140830174914 Git commit info: 2014-08-30 06:16:06 83cb8148
(In reply to Pi Wei Cheng [:piwei] from comment #11) > This issue does NOT repro on Flame 2.1, Flame 2.0 > > After reboot and with Settings app open, running 'adb shell top -m 5' > returns the following result: > > On 2.1 Flame: > > User 28%, System 26%, IOW 0%, IRQ 0% > User 181 + Nice 15 + Sys 185 + Idle 305 + IOW 3 + IRQ 0 + SIRQ 4 = 693 > > PID PR CPU% S #THR VSS RSS PCY UID Name > 295 1 30% S 60 212796K 91740K root /system/b2g/b2g > 1586 1 3% S 18 83072K 34072K u0_a1586 /system/b2g/b2g > 285 0 2% S 7 23804K 7508K fg media /system/bin/mediaserver > 1333 1 2% S 19 70548K 27636K u0_a1333 /system/b2g/b2g > 96 1 1% S 1 0K 0K root ksmd > This obviously shows that the issues does reproduce on Flame running master builds !
Pi Wei, your comment 11 does not makes sense: the figures you expose shows that we have a 30% CPU usage while doing nothing on master, and that it was not the case for previous versions. I suspect you wanted to say that the issue DOES reproduce :). Given the PID, I suspect it's the main B2G process.
Flags: needinfo?(pcheng)
I saw the similar issue(CPU usage 25~50%) on my Flame with JB debug build(B2G_DEBUG=1) from m-c. But it seems to happening when I use release build. 130|root@flame:/ # top -m 5 -t User 48%, System 2%, IOW 0%, IRQ 0% User 240 + Nice 62 + Sys 15 + Idle 302 + IOW 0 + IRQ 0 + SIRQ 0 = 619 PID TID PR CPU% S VSS RSS PCY UID Thread Proc 12569 12569 1 32% R 215244K 99264K root b2g /system/b2g/b2g 12801 12801 1 10% S 134596K 55636K u0_a1280 Homescreen /system/b2g/plugin-container 12569 12597 0 6% S 215244K 99264K root DOM Worker /system/b2g/b2g 13995 13995 0 1% R 1276K 584K root top top 96 96 1 0% S 0K 0K root ksmd
I noticed that limiting the memory and the number of CPUs on Flame does help to reproduce the issue.
s/it seems to happening/it seems not happening/
(In reply to Alexandre LISSY :gerard-majax from comment #18) > Pi Wei, your comment 11 does not makes sense: the figures you expose shows > that we have a 30% CPU usage while doing nothing on master, and that it was > not the case for previous versions. I suspect you wanted to say that the > issue DOES reproduce :). Given the PID, I suspect it's the main B2G process. Yes and no. CPU usage was NOT as high as what was shown at comment 0, comment 4, or title of this bug. The regression within Flame device itself is a bug, but I'm not sure if it's the same cause as this one. Flagging our test lead to make a decision.
Flags: needinfo?(pcheng) → needinfo?(jmitchell)
I think we are just lacking communication on what should be considered a 'repro' on this. The numbers Pi Wei posts are significantly different than the description (90-95%). Has anyone determined what the standard (acceptable) numbers are (just idling on homescreen)?
Flags: needinfo?(jmitchell)
(In reply to Joshua Mitchell [:Joshua_M] from comment #23) > I think we are just lacking communication on what should be considered a > 'repro' on this. The numbers Pi Wei posts are significantly different than > the description (90-95%). Has anyone determined what the standard > (acceptable) numbers are (just idling on homescreen)? I don't think that we have any problem with communication here. Please read all comments carefully. Moreover qa made test on different device which is much stronger than hamachi. As gerard-majax wrote in Comment 20 'limiting the memory and the number of CPUs on Flame does help' because on stronger device problem can be easily missed. In Comment 11 qa pasted top output which clearly shown that there is a load on 2.1 but smaller - as I guess becuse Flame has two Cortex A7 1,3GHz cores when hamachi has one weak Cortex A5 1GHz core. Finally guys are fighting now to solve it then the issue exists and looks serious.
We cannot exclude the possibility that these could be different issues, as the original problem requires to run an app to reproduce, while other testing results show high CPU load when idling on homescreen. Shall we file a different bug for tracking, or modify the title if decided to track them as same high CPU loading issue?
(In reply to Shian-Yow Wu [:swu] from comment #25) > We cannot exclude the possibility that these could be different issues, as > the original problem requires to run an app to reproduce, while other > testing results show high CPU load when idling on homescreen. > Shall we file a different bug for tracking, or modify the title if decided > to track them as same high CPU loading issue? Feel free to change title when we collected already much information here. From my experience it looks like same issue because, read carefully comments, load fall from ~95% to ~70% (in comment 10) and next to the current level (in comment 14). In case of 95% load again I will create a new bug and like to this one.
A similar issue on Flame device is being dealt with in bug 1062119.
(In reply to Shian-Yow Wu [:swu] from comment #19) > I saw the similar issue(CPU usage 25~50%) on my Flame with JB debug > build(B2G_DEBUG=1) from m-c. But it seems to happening when I use release > build. > > 130|root@flame:/ # top -m 5 -t > > > User 48%, System 2%, IOW 0%, IRQ 0% > User 240 + Nice 62 + Sys 15 + Idle 302 + IOW 0 + IRQ 0 + SIRQ 0 = 619 > > PID TID PR CPU% S VSS RSS PCY UID Thread Proc > 12569 12569 1 32% R 215244K 99264K root b2g > /system/b2g/b2g > 12801 12801 1 10% S 134596K 55636K u0_a1280 Homescreen > /system/b2g/plugin-container > 12569 12597 0 6% S 215244K 99264K root DOM Worker > /system/b2g/b2g > 13995 13995 0 1% R 1276K 584K root top top > 96 96 1 0% S 0K 0K root ksmd I tested it again on Flame device with B2G_DEBUG=1 by today's gecko/gaia code, and the issue in comment 19 which idles in homescreen is not reproducible. Gecko: 6b8da5940f74 Gaia: b630b8bcaf9653885539d4449bc65c3b592bd752
QA Whiteboard: [QAnalyst-Triage+] → [QAnalyst-Triage+][lead-review+]
After restart and opened a few webpages I've got again same as in https://bugzilla.mozilla.org/show_bug.cgi?id=1062255#c41, the new profile: http://people.mozilla.org/~bgirard/cleopatra/#report=30441400a456e1d2a6fb6aff7349a0fcda7b0338 taken on: Alcatel One Touch Fire production (got from T-mobile Poland) B2G version: 2.2.0.0-prerelease master Platform version: 35.0a1 Build Identifier: 20140914040220 Git commit info: 2014-09-13 09:23:34 e5da0e46
Flags: needinfo?(lissyx+mozillians)
Kevin and Dale, you may want to have a look at the profile provided in comment 29. We are not yet sure if it's reproduced with a fresh profile, but there's something that goes bad in places.js :(
Flags: needinfo?(lissyx+mozillians)
Flags: needinfo?(kgrandon)
Flags: needinfo?(dale)
This seems expected for opening webpages. I suppose we should compare this to the 2.0 browser as that release was also storing visits. I wonder if the use of datastore here is causing some additional latency. Leaving ni? to investigate further.
Im gonna leave the needinfo on me, is there any way to read that profile to say how many times 'edit place' was called, also mac could you give me a brief description of what you did while generating that profile? addVisit when called on a url that is counted in the top sites (top 6 most visited) is going to generate a screenshot which is pretty cpu intensive although I dont see that in the profile, we may want to cut down how many times we take that screenshot Also we have a global lock on editPlace which ensures we dont lose any data, since the datastore api is fairly poor it means we have a lock on the entire store, we may want to do a pessimistic lock on each individual place instead if it does end up spinning a loop, it will always eventually end, but if we manage to fire a whole bunch of events at the same time (which is pretty hard) then it could end up looping a bit much
Something that may also help here is only saving the place on mozbrowserloadend. We can build a local cache map of all of the data/icons, so we only hit IDB/datastore a single time. I can look at working on this if needed.
Take for example this dump of events that we get for loading mozilla.org: Event, url: applocationchange https://www.mozilla.org/en-US/ Event, url: apptitlechange https://www.mozilla.org/en-US/ Event, url: appiconchange https://www.mozilla.org/en-US/ Event, url: appiconchange https://www.mozilla.org/en-US/ Event, url: apploaded https://www.mozilla.org/en-US/ Right now it appears that we would both fetch and get the place from datastore for each event. My suggestion is that we only do this once per website if possibe. Probably after some timeout after applocationchange, or on apploaded.
Flags: needinfo?(kgrandon)
(In reply to Dale Harvey (:daleharvey) from comment #32) > ... also mac could you give me a brief description of what you did while generating that profile? I'm not sure if I clearly understood your question but profile was taken when device was idling - on the Homescreen and none apps was running (empty task manager).
Depends on: 1068888
(In reply to mac from comment #35) > (In reply to Dale Harvey (:daleharvey) from comment #32) > > ... also mac could you give me a brief description of what you did while generating that profile? > > I'm not sure if I clearly understood your question but profile was taken > when device was idling - on the Homescreen and none apps was running (empty > task manager). But the profile in comment 29 shows browser activity correct? Are these the same issues? In any case I've filed bug 1068888 to spin off the places work, because at this point I'm not sure if it's 100% related.
Kevins got a patch in https://bugzilla.mozilla.org/show_bug.cgi?id=1068888 which should lower the amount places is hit so clearing needinfo here
Flags: needinfo?(dale)
Issue still exists on: Alcatel One Touch Fire production (got from T-mobile Poland) B2G version: 2.2.0.0-prerelease master Platform version: 36.0a1 Build Identifier: 201411180220608 Git commit info: 2014-11-18 19:09:06 3cad1e7b The profile: http://people.mozilla.org/~bgirard/cleopatra/#report=30441400a456e1d2a6fb6aff7349a0fcda7b0338
Bulk edit to clear old and out of date needinfo requests that I never responded to. I'm assuming that these are no longer relevant. If you are still waiting for an answer from me, please set needinfo? again.
Flags: needinfo?(dflanagan)
Firefox OS is not being worked on
Status: UNCONFIRMED → RESOLVED
Closed: 8 years ago
Resolution: --- → WONTFIX
You need to log in before you can comment on or make changes to this bug.

Attachment

General

Creator:
Created:
Updated:
Size: