Code coverage ingestion is broken since 14/09/2024 (last working revision 99b3ca864422ae96e52d061a149c9e454a986443)
Categories
(Testing :: Code Coverage, defect)
Tracking
(firefox133 wontfix)
| Tracking | Status | |
|---|---|---|
| firefox133 | --- | wontfix |
People
(Reporter: marco, Unassigned)
References
Details
(Keywords: regression)
Attachments
(2 files)
We are seeing:
Exception:
grcovfailed with code: -9.
in ingestion tasks.
| Reporter | ||
Updated•1 year ago
|
| Reporter | ||
Comment 1•1 year ago
|
||
The first failing revision is a51b9d3e7251d81217a4d7b3622a1658e9a23dfc.
Comment 2•1 year ago
|
||
This bug has been marked as a regression. Setting status flag for Nightly to affected.
Comment 3•1 year ago
|
||
The severity field is not set for this bug.
:marco, could you have a look please?
For more information, please visit BugBot documentation.
| Reporter | ||
Updated•1 year ago
|
| Reporter | ||
Comment 4•1 year ago
|
||
It's still unclear what caused this. We tried increasing RAM on the machine, but that didn't help. The errors in the grcov logs are all just warnings.
Bug 1925875, which we noticed recently, is likely pre-existing and also doesn't cause grcov to completely fail.
Updated•1 year ago
|
Comment 5•1 year ago
|
||
Updated•1 year ago
|
Comment 6•1 year ago
|
||
Authored by https://github.com/jcristau
https://github.com/mozilla-releng/fxci-config/commit/7910af85780466080594a328115f08a3e8e94d63
[main] Bug 1925873 - bump code-coverage/bot-gcp pool to c2-standard-30 (#394)
Comment 7•1 year ago
|
||
This didn't help, grcov still goes oom even with 120G.
systemd-udevd invoked oom-killer: gfp_mask=0x100cca(GFP_HIGHUSER_MOVABLE), order=0, oom_score_adj=0
CPU: 17 PID: 7067 Comm: systemd-udevd Kdump: loaded Tainted: G OE 5.4.0-1106-gcp #115~18.04.1-Ubuntu
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 02/12/2025
Call Trace:
dump_stack+0x6d/0x8b
dump_header+0x4f/0x200
oom_kill_process+0xec/0x140
out_of_memory+0x117/0x570
__alloc_pages_slowpath+0xada/0xec0
__alloc_pages_nodemask+0x2cd/0x320
alloc_pages_current+0x6a/0xe0
__page_cache_alloc+0x6a/0xa0
pagecache_get_page+0xab/0x2c0
filemap_fault+0x685/0xb80
? unlock_page_memcg+0x12/0x20
? page_add_file_rmap+0x13a/0x180
? xas_load+0xc/0x80
? xas_find+0x16f/0x1b0
? filemap_map_pages+0x181/0x3b0
ext4_filemap_fault+0x31/0x50
__do_fault+0x57/0x158
__handle_mm_fault+0xdae/0x1240
? __seccomp_filter+0x7d/0x710
handle_mm_fault+0xcb/0x210
__do_page_fault+0x2a1/0x4d0
do_page_fault+0x2c/0xe0
page_fault+0x34/0x40
RIP: 0033:0x7f40dd1d1ad5
Code: Bad RIP value.
RSP: 002b:00007fffa50c24e8 EFLAGS: 00010216
RAX: 00007fffa50c24f0 RBX: 00007fffa50c3615 RCX: 0000000000000001
RDX: 000000000000000f RSI: 00005618f3c94de1 RDI: 00007fffa50c24f0
RBP: 00007fffa50c35c0 R08: 00005618f55b55a0 R09: 00007f40dd505c40
R10: 0000000000000007 R11: 00007f40dd2c9494 R12: 00007fffa50c3611
R13: 0000000000000001 R14: 00005618f55b55a8 R15: 00007fffa50c2520
Mem-Info:
active_anon:30388970 inactive_anon:67 isolated_anon:0
active_file:6082 inactive_file:6082 isolated_file:12
unevictable:0 dirty:2 writeback:0 unstable:0
slab_reclaimable:41375 slab_unreclaimable:59160
mapped:141 shmem:548 pagetables:83498 bounce:0
free:138523 free_pcp:0 free_cma:0
Node 0 active_anon:121555880kB inactive_anon:268kB active_file:24328kB inactive_file:24328kB unevictable:0kB isolated(anon):0kB isolated(file):48kB mapped:564kB dirty:8kB writeback:0kB shmem:2192kB shmem_thp: 0kB shmem_pmdmapped: 0kB anon_thp: 0kB writeback_tmp:0kB unstable:0kB all_unreclaimable? no
Node 0 DMA free:15584kB min:8kB low:20kB high:32kB active_anon:0kB inactive_anon:0kB active_file:0kB inactive_file:0kB unevictable:0kB writepending:0kB present:15920kB managed:15584kB mlocked:0kB kernel_stack:0kB pagetables:0kB bounce:0kB free_pcp:0kB local_pcp:0kB free_cma:0kB
lowmem_reserve[]: 0 2759 120508 120508 120508
Node 0 DMA32 free:472508kB min:1544kB low:4368kB high:7192kB active_anon:2386388kB inactive_anon:0kB active_file:0kB inactive_file:0kB unevictable:0kB writepending:0kB present:3126072kB managed:2863888kB mlocked:0kB kernel_stack:0kB pagetables:4656kB bounce:0kB free_pcp:0kB local_pcp:0kB free_cma:0kB
lowmem_reserve[]: 0 0 117748 117748 117748
Node 0 Normal free:66000kB min:66024kB low:186596kB high:307168kB active_anon:119169552kB inactive_anon:268kB active_file:24408kB inactive_file:23820kB unevictable:0kB writepending:0kB present:122683392kB managed:120581724kB mlocked:0kB kernel_stack:10224kB pagetables:329336kB bounce:0kB free_pcp:0kB local_pcp:0kB free_cma:0kB
lowmem_reserve[]: 0 0 0 0 0
Node 0 DMA: 0*4kB 0*8kB 0*16kB 1*32kB (U) 1*64kB (U) 1*128kB (U) 0*256kB 0*512kB 1*1024kB (U) 1*2048kB (M) 3*4096kB (M) = 15584kB
Node 0 DMA32: 3*4kB (UM) 8*8kB (UM) 5*16kB (M) 1*32kB (M) 1*64kB (U) 2*128kB (UM) 1*256kB (U) 0*512kB 1*1024kB (U) 2*2048kB (UM) 114*4096kB (M) = 472828kB
Node 0 Normal: 5305*4kB (UME) 2013*8kB (UME) 478*16kB (UME) 732*32kB (UME) 0*64kB 0*128kB 0*256kB 0*512kB 0*1024kB 0*2048kB 0*4096kB = 68396kB
Node 0 hugepages_total=0 hugepages_free=0 hugepages_surp=0 hugepages_size=1048576kB
Node 0 hugepages_total=0 hugepages_free=0 hugepages_surp=0 hugepages_size=2048kB
12775 total pagecache pages
0 pages in swap cache
Swap cache stats: add 0, delete 0, find 0/0
Free swap = 0kB
Total swap = 0kB
31456346 pages RAM
0 pages HighMem/MovableOnly
591047 pages reserved
0 pages cma reserved
0 pages hwpoisoned
Tasks state (memory values in pages):
[ pid ] uid tgid total_vm rss pgtables_bytes swapents oom_score_adj name
[ 720] 0 720 51110 453 421888 0 0 systemd-journal
[ 742] 0 742 26477 65 98304 0 0 lvmetad
[ 758] 0 758 11572 442 122880 0 -1000 systemd-udevd
[ 1741] 0 1741 1129 17 57344 0 0 none
[ 2000] 100 2000 20012 165 188416 0 0 systemd-network
[ 2035] 101 2035 17656 152 176128 0 0 systemd-resolve
[ 2228] 0 2228 42815 2049 217088 0 0 networkd-dispat
[ 2264] 102 2264 93266 7167 294912 0 0 rsyslogd
[ 2275] 0 2275 554624 1604 356352 0 0 google_osconfig
[ 2296] 0 2296 40271 33 86016 0 0 lxcfs
[ 2354] 103 2354 12526 176 139264 0 -900 dbus-daemon
[ 2394] 0 2394 72000 234 196608 0 0 accounts-daemon
[ 2435] 0 2435 7084 52 98304 0 0 atd
[ 2474] 0 2474 1159 17 53248 0 0 sshguard-journa
[ 2475] 0 2475 562469 3341 421888 0 -999 containerd
[ 2476] 0 2476 52264 119 364544 0 0 journalctl
[ 2477] 0 2477 4093 99 61440 0 0 sshguard
[ 2523] 0 2523 46922 1972 262144 0 0 unattended-upgr
[ 2539] 0 2539 497315 1603 307200 0 -999 google_guest_ag
[ 2552] 0 2552 4105 37 73728 0 0 agetty
[ 2654] 0 2654 3724 32 69632 0 0 agetty
[ 2657] 0 2657 1159 16 57344 0 0 sshg-fw
[ 2674] 0 2674 72221 197 196608 0 0 polkitd
[ 3438] 0 3438 18076 183 180224 0 -1000 sshd
[ 3446] 0 3446 15503 145 159744 0 0 systemd-logind
[ 3462] 0 3462 7938 72 102400 0 0 cron
[ 3783] 111 3783 25296 54 102400 0 0 chronyd
[ 4252] 0 4252 868678 3235 466944 0 -900 snapd
[ 5126] 0 5126 931112 7101 770048 0 -500 dockerd
[ 5315] 0 5315 3330 59 69632 0 0 start-docker-wo
[ 5325] 0 5325 308519 1106 118784 0 -1000 start-worker
[ 5326] 0 5326 6632 42 94208 0 0 logger
[ 5345] 0 5345 236360 14547 2428928 0 -1000 node
[ 5502] 0 5502 180125 527 110592 0 -998 containerd-shim
[ 5524] 0 5524 27818 889 90112 0 0 taskcluster-pro
[ 5610] 0 5610 309021 257 167936 0 -500 docker-proxy
[ 5618] 0 5618 272027 258 139264 0 -500 docker-proxy
[ 5632] 0 5632 180125 528 118784 0 -998 containerd-shim
[ 5654] 0 5654 27655 790 81920 0 0 livelog
[ 5782] 0 5782 180125 605 114688 0 -998 containerd-shim
[ 5805] 0 5805 585292 40591 516096 0 0 code-coverage-r
[ 6867] 0 6867 50740814 30292854 334737408 0 0 grcov
[ 7067] 0 7067 11572 443 110592 0 0 systemd-udevd
oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=/,mems_allowed=0,global_oom,task_memcg=/docker/2cf03290442e3601f7207f04585174c3e39ef86dbaa317841156f957a10c438a,task=grcov,pid=6867,uid=0
Out of memory: Killed process 6867 (grcov) total-vm:202963256kB, anon-rss:121171412kB, file-rss:4kB, shmem-rss:0kB, UID:0 pgtables:326892kB oom_score_adj:0
oom_reaper: reaped process 6867 (grcov), now anon-rss:0kB, file-rss:0kB, shmem-rss:0kB
| Reporter | ||
Updated•1 year ago
|
| Reporter | ||
Comment 9•1 year ago
|
||
(In reply to Marco Castelluccio [:marco] from comment #4)
It's still unclear what caused this. We tried increasing RAM on the machine, but that didn't help. The errors in the grcov logs are all just warnings.
Bug 1925875, which we noticed recently, is likely pre-existing and also doesn't cause grcov to completely fail.
After more investigation, bug 1925875 might somehow be related.
With one of the malformed files, grcov 0.7 is running OOM, while grcov 0.9 is not (however, it returns an error).
After fixing obvious mistakes in the file (records cut too early), grcov 0.7 is still running OOM, while grcov 0.9 is not (without returning an error).
We can't simply switch to grcov 0.9 yet because it would return an error and not produce coverage data, we will need to fix https://github.com/mozilla/grcov/issues/1242 first.
Comment 10•1 year ago
|
||
Comment 11•1 year ago
|
||
Authored by https://github.com/jcristau
https://github.com/mozilla-releng/fxci-config/commit/3cf70452c0990bb9731774ef9a4dce56e05c1f2a
[main] Bug 1925873 - switch code-coverage/bot-gcp pool back to c2-standard-4 (#396)
| Reporter | ||
Comment 12•11 months ago
|
||
This was fixed thanks to a grcov update.
| Reporter | ||
Comment 13•11 months ago
|
||
Full story for posterity:
The OOM was due to malformed data which triggered abnormal memory usage in grcov. The latest grcov version didn't have this bug, but just exited with an error when it encountered malformed data.
So we made it so grcov doesn't exit with an error in case of malformed data.
It was then failing because of a OOM because the latest grcov version was built normally without tcmalloc, which we used to use because it reduced memory usage by a lot.
So I made it use tcmalloc again.
Updated•11 months ago
|
Description
•