Open Bug 1908798 Opened 2 years ago Updated 20 hours ago

Crash in [@ IPCError-browser | GPUProcessKill]

Categories

(Core :: Graphics: WebRender, defect, P2)

Unspecified
Windows 10
defect

Tracking

()

Tracking Status
firefox-esr115 --- unaffected
firefox-esr128 --- affected
firefox128 --- wontfix
firefox129 --- wontfix
firefox130 --- wontfix
firefox131 --- wontfix
firefox132 --- wontfix
firefox133 --- wontfix
firefox134 --- wontfix

People

(Reporter: mccr8, Assigned: sotaro, NeedInfo)

References

(Depends on 1 open bug, Regression)

Details

(Keywords: crash, regression, topcrash, Whiteboard: [tbird crash])

Crash Data

Crash report: https://crash-stats.mozilla.org/report/index/c95a1ffd-5281-48ee-96a4-59e440240718

Reason: EXCEPTION_BREAKPOINT

Top 10 frames:

0  win32u.dll  ZwUserMsgWaitForMultipleObjectsEx
1  user32.dll  RealMsgWaitForMultipleObjectsEx(unsigned long, void* const*, unsigned long, u...
2  xul.dll  mozilla::widget::WinUtils::WaitForMessage(unsigned long)  widget/windows/WinUtils.cpp:454
3  xul.dll  nsAppShell::ProcessNextNativeEvent(bool)  widget/windows/nsAppShell.cpp:796
4  xul.dll  nsBaseAppShell::DoProcessNextNativeEvent(bool)  widget/nsBaseAppShell.cpp:131
4  xul.dll  nsBaseAppShell::OnProcessNextEvent(nsIThreadInternal*, bool)  widget/nsBaseAppShell.cpp:267
4  xul.dll  nsThread::ProcessNextEvent(bool, bool*)  xpcom/threads/nsThread.cpp:1119
4  xul.dll  NS_ProcessNextEvent(nsIThread*, bool)  xpcom/threads/nsThreadUtils.cpp:480
5  xul.dll  mozilla::ipc::MessagePump::Run(base::MessagePump::Delegate*)  ipc/glue/MessagePump.cpp:107
6  xul.dll  MessageLoop::RunInternal()  ipc/chromium/src/base/message_loop.cc:370

I don't know if this is the same as bug 1900134 or not, but the signature is slightly different, so I'll file a new bug for it. Decent volume on Nightly.

The bug is linked to a topcrash signature, which matches the following criteria:

  • Top 10 desktop browser crashes on nightly (startup)
  • Top 5 GPU process crashes on release (startup)

For more information, please visit BugBot documentation.

This is going to be hard to investigate unless we can get a reproduction case or a more informative call stack.

The startup crashes I checked have GraphicsCriticalError of shader-cache: Shader disk cache is not supported. The non-startup crashes seem to have timeouts of different kinds, including shader-cache: Timed out before finishing loads and Killing GPU process due to IPC reply timeout.

Jamie, you're in that code that sets those errors. Are they signifying anything useful in this case?

Flags: needinfo?(jnicol)

I'm not sure why the signature for this and bug bug 1900134 differ, but the other bug is all android crashes and this is windows, so keeping them separate makes sense.

As I said in the other bug, this crash was introduced by bug 1880503, but these are pre-existing issues - previously we would just silently kill the GPU process, whereas now we get a crash report too.

I think the significance of the shader cache errors is that they indicate a really old or slow device. These crashes mostly seem to be occurring during webrender initialization (Compositor thread is in WebRenderAPI::Create, renderer thread at various stages of initialization.) Perhaps we need a longer timeout than 10 seconds? I expect the current experience is pretty terrible for affected users - several 10 second timeouts attempting to initialize webrender then eventually the GPU process gets disabled, then presumably a >10s wait to initialize webrender in the parent process.

Flags: needinfo?(jnicol)
Severity: -- → S2
Priority: -- → P2

Dropping this into WebRender for a look-see.

Glenn, not expecting you'll be able to do a lot with this, but given the low crash volume now, do you think this remains valid as an S2?

Component: Graphics → Graphics: WebRender
Flags: needinfo?(gwatson)

It does seem to have spiked and be ~100 / day, S2 is probably reasonable if we can do something about it.

Flags: needinfo?(gwatson)

Adding a responsible party, and setting to P1 so it shows up on my internal tracking radar.

Assignee: nobody → gwatson
Priority: P2 → P1

Sotaro, in the crash report it looks like the Compositor thread is trying to shut down - any ideas what might cause this IPC error or what we should look at?

Flags: needinfo?(sotaro.ikeda.g)
Flags: needinfo?(sotaro.ikeda.g)
Keywords: regression
Regressed by: 1880503

As in comment 3, the crash is pre-existing issues for long time. Just bug 1880503 has just been changed to generate a crash report.

The crash happened by killing GPU process from CompositorManagerChild::ShouldContinueFromReplyTimeout() by calling GPUProcessManager::KillProcess().

CompositorManagerChild::ShouldContinueFromReplyTimeout() is called when sync IPC failed with reply timeout from MessageChannel::ShouldContinueFromTimeout().

Reply timeout is set from CompositorManagerChild::SetReplyTimeout().

Set release status flags based on info from the regressing bug 1880503

From the crash reports, majority of the cashes seemed to happen within 60 sec from starting Firefox.

In this case, the crash reports tend to have the following functions in calling stacks.

There are several sync IPC message under gfx.
https://searchfox.org/mozilla-central/search?q=+sync+&path=gfx**.ipdl&case=false&regexp=false

The crash of comment 10 might be triggered by PWebRenderBridgeChild::SendEnsureConnected(). It is sync IPC.

It might be possible to relax CompositorManagerChild::ShouldContinueFromReplyTimeout() until RenderThread::InitDeviceTask() complete in GPU process.

As in comment 3, the crash seemed to tend to happen with a really old or slow device. With these devices, there might be cases that Windows system call became very slow. It might trigger sync IPC reply timeout.

From the following document, I obtained Firefox profiler during Firefox startup, and synchronization IPC did not take a long time during normal Firefox startup.

https://profiler.firefox.com/docs/#/./guide-startup-shutdown

Set release status flags based on info from the regressing bug 1880503

A friend sent me a crash report which came to this signature. He says he "has been plagued by them for the last couple months". The user experiences for this was the following:

  • Had long lived firefox session. Opened a link, and firefox appeared to freeze.
  • Windows that were minimized wouldn't come back
  • Only process in task manager not responding still seemed to be allocating memory (50kb/s -- not sure how he measured this)
  • Ended up killing process.

If there's other debugging information desired, I can relay instructions

Sotaro, I'm assigning this to you as you seem to have an understanding of what's going on here. Do we have a path forward to work on this?

Assignee: gwatson → sotaro.ikeda.g

(In reply to Matthew Gaudet (he/him) [:mgaudet] from comment #16)

A friend sent me a crash report which came to this signature. He says he "has been plagued by them for the last couple months". The user experiences for this was the following:

In the crash report, Compositor thread called CompositorBridgeParent::PauseComposition(). And RendererThread called
SyncObjectD3D11Host::Synchronize().

(In reply to Sotaro Ikeda [:sotaro] from comment #11)

There are several sync IPC message under gfx.
https://searchfox.org/mozilla-central/search?q=+sync+&path=gfx**.ipdl&case=false&regexp=false

The crash of comment 10 might be triggered by PWebRenderBridgeChild::SendEnsureConnected(). It is sync IPC.

Since there are a lot of sync IPC exist, it seems better to use a longer timeout than 10 seconds like comment 3 at first.

Depends on: 1922157

(In reply to Sotaro Ikeda [:sotaro] from comment #18)

(In reply to Matthew Gaudet (he/him) [:mgaudet] from comment #16)

A friend sent me a crash report which came to this signature. He says he "has been plagued by them for the last couple months". The user experiences for this was the following:

In the crash report, Compositor thread called CompositorBridgeParent::PauseComposition(). And RendererThread called
SyncObjectD3D11Host::Synchronize().

CompositorBridgeParent::PauseComposition() is called like the following sequence.

CanonicalBrowsingContext::RecomputeAppWindowVisibility()
->nsIWidget::PauseOrResumeCompositor()
->CompositorBridgeChild::SendPause()
// sync IPC to GPU process
->CompositorBridgeParent::RecvPause()
->CompositorBridgeParent::PauseComposition()
->WebRenderBridgeParent::Pause()
->WebRenderAPI::Pause()
->//Post task to RenderThread
->RenderThread::Pause()
->RenderCompositorANGLE::Pause() {} //Do nothing

WebRenderAPI::Pause() post task to RenderThread by calling WebRenderAPI::RunOnRenderThread().

It end up to call "self.low_priority_scene_sender.send(msg).unwrap()" in RenderApi::send_external_event(). It might take long time until the posted task run in RenderThread, since low_priority_scene_sender is used.

It is nice if RenderThread::Pause() could be called without low_priority_scene_sender. In this case, we need to care about consistency between WebRender and GL/EGL(EGLSurface).

Short term workaround for Windows is skip RenderThread::Pause() and RenderThread::Resume(), since RenderCompositorANGLE::Pause() and RenderCompositorANGLE::Resume() do nothing.

Depends on: 1922214

Bug 1922214 was created for comment 22.

There were several crashes with CompositorBridgeParent::RecvFlushRendering() in Compositor thread. It is triggered like the following

nsViewManager::Refresh()
->WebRenderLayerManager::FlushRendering()
->CompositorBridgeChild::SendFlushRendering() // Triggers sync IPC
->// sync IPC
->CompositorBridgeParent::RecvFlushRendering()

On Windows, WebRenderLayerManager::FlushRendering() always calls CompositorBridgeChild::SendFlushRendering(), since pref layers.force-synchronous-resize = true.
https://searchfox.org/mozilla-central/rev/ce404cd26e52d09e6a48d664c1986da25df50484/modules/libpref/init/StaticPrefList.yaml#8543

It is nice if we could remove/reduce sync FlushRendering on Windows.

Depends on: 1922721
Depends on: 1923263

Since Bug 1922157 fix, the amount of crashes seems to have decreased in nightly.

Severity: S2 → S3
Priority: P1 → P2
Whiteboard: [tbird crash]

Sotaro, it looks like the crash volume has increased again. Is there anything blocking us from moving forward with a fix here?

Flags: needinfo?(sotaro.ikeda.g)

Two Crashes with this signature.

  • Firefox 141.0b3
  • Firefox 142.0a1

Firefox 141.0b3 Crash Report [@ IPCError-browser | GPUProcessKill ]
Crash ID: bp-3180fcf3-bfb3-474e-ab83-13d8f0250815

Frame 	Module 	Signature 	Source 	Trust
0 	win32u.dll 	ZwUserMsgWaitForMultipleObjectsEx 		context
1 	user32.dll 	RealMsgWaitForMultipleObjectsEx(unsigned long, void* const*, unsigned long, unsigned long, unsigned long) 		cfi 
2 	xul.dll 	mozilla::widget::WinUtils::WaitForMessage(unsigned long) 	widget/windows/WinUtils.cpp:451 	cfi
3 	xul.dll 	nsAppShell::ProcessNextNativeEvent(bool) 	widget/windows/nsAppShell.cpp:796 	cfi
4 	xul.dll 	nsBaseAppShell::DoProcessNextNativeEvent(bool) 	widget/nsBaseAppShell.cpp:131 	inlined
4 	xul.dll 	nsBaseAppShell::OnProcessNextEvent(nsIThreadInternal*, bool) 	widget/nsBaseAppShell.cpp:267 	inlined
4 	xul.dll 	nsThread::ProcessNextEvent(bool, bool*) 	xpcom/threads/nsThread.cpp:1098 	inlined
4 	xul.dll 	NS_ProcessNextEvent(nsIThread*, bool) 	xpcom/threads/nsThreadUtils.cpp:480 	cfi
5 	xul.dll 	mozilla::ipc::MessagePump::Run(base::MessagePump::Delegate*) 	ipc/glue/MessagePump.cpp:107 	cfi
6 	xul.dll 	MessageLoop::RunInternal() 	ipc/chromium/src/base/message_loop.cc:369 	inlined
6 	xul.dll 	MessageLoop::RunHandler() 	ipc/chromium/src/base/message_loop.cc:362 	cfi
7 	xul.dll 	MessageLoop::Run() 	ipc/chromium/src/base/message_loop.cc:344 	inlined
7 	xul.dll 	nsBaseAppShell::Run() 	widget/nsBaseAppShell.cpp:148 	cfi
8 	xul.dll 	nsAppShell::Run() 	widget/windows/nsAppShell.cpp:673 	cfi
9 	xul.dll 	XRE_RunAppShell() 	toolkit/xre/nsEmbedFunctions.cpp:652 	inlined
9 	xul.dll 	mozilla::ipc::MessagePumpForChildProcess::Run(base::MessagePump::Delegate*) 	ipc/glue/MessagePump.cpp:235 	cfi
10 	xul.dll 	MessageLoop::RunInternal() 	ipc/chromium/src/base/message_loop.cc:369 	inlined
10 	xul.dll 	MessageLoop::RunHandler() 	ipc/chromium/src/base/message_loop.cc:362 	cfi 
11 	xul.dll 	MessageLoop::Run() 	ipc/chromium/src/base/message_loop.cc:344 	inlined
11 	xul.dll 	XRE_InitChildProcess(int, char**, XREChildData const*) 	toolkit/xre/nsEmbedFunctions.cpp:590 	inlined
11 	xul.dll 	mozilla::BootstrapImpl::XRE_InitChildProcess(int, char**, XREChildData const*) 	toolkit/xre/Bootstrap.cpp:60 	cfi
12 	firefox.exe 	NS_internal_main(int, char**, char**) 	browser/app/nsBrowserApp.cpp:397 	inlined
12 	firefox.exe 	wmain(int, wchar_t**) 	toolkit/xre/nsWindowsWMain.cpp:151 	cfi
13 	firefox.exe 	invoke_main() 	/builds/worker/workspace/obj-build/browser/app/D:/a/_work/1/s/src/vctools/crt/vcstartup/src/startup/exe_common.inl:90 	inlined
13 	firefox.exe 	__scrt_common_main_seh() 	/builds/worker/workspace/obj-build/browser/app/D:/a/_work/1/s/src/vctools/crt/vcstartup/src/startup/exe_common.inl:288 	cfi
14 	kernel32.dll 	BaseThreadInitThunk 		cfi
15 	ntdll.dll 	RtlUserThreadStart 		cfi

Firefox 142.0a1 Crash Report [@ IPCError-browser | GPUProcessKill ]
Crash ID: bp-4107237a-3317-4541-9e81-259ac0250815

Frame 	Module 	Signature 	Source 	Trust
0 	win32u.dll 	ZwUserMsgWaitForMultipleObjectsEx 		context
1 	user32.dll 	RealMsgWaitForMultipleObjectsEx(unsigned long, void* const*, unsigned long, unsigned long, unsigned long) 		cfi
2 	xul.dll 	mozilla::widget::WinUtils::WaitForMessage(unsigned long) 	widget/windows/WinUtils.cpp:451 	cfi
3 	xul.dll 	nsAppShell::ProcessNextNativeEvent(bool) 	widget/windows/nsAppShell.cpp:796 	cfi
4 	xul.dll 	nsBaseAppShell::DoProcessNextNativeEvent(bool) 	widget/nsBaseAppShell.cpp:131 	inlined
4 	xul.dll 	nsBaseAppShell::OnProcessNextEvent(nsIThreadInternal*, bool) 	widget/nsBaseAppShell.cpp:267 	inlined
4 	xul.dll 	nsThread::ProcessNextEvent(bool, bool*) 	xpcom/threads/nsThread.cpp:1098 	inlined
4 	xul.dll 	NS_ProcessNextEvent(nsIThread*, bool) 	xpcom/threads/nsThreadUtils.cpp:480 	cfi
5 	xul.dll 	mozilla::ipc::MessagePump::Run(base::MessagePump::Delegate*) 	ipc/glue/MessagePump.cpp:107 	cfi
6 	xul.dll 	MessageLoop::RunInternal() 	ipc/chromium/src/base/message_loop.cc:369 	inlined
6 	xul.dll 	MessageLoop::RunHandler() 	ipc/chromium/src/base/message_loop.cc:362 	cfi
7 	xul.dll 	MessageLoop::Run() 	ipc/chromium/src/base/message_loop.cc:344 	inlined
7 	xul.dll 	nsBaseAppShell::Run() 	widget/nsBaseAppShell.cpp:148 	cfi
8 	xul.dll 	nsAppShell::Run() 	widget/windows/nsAppShell.cpp:673 	cfi
9 	xul.dll 	XRE_RunAppShell() 	toolkit/xre/nsEmbedFunctions.cpp:652 	inlined
9 	xul.dll 	mozilla::ipc::MessagePumpForChildProcess::Run(base::MessagePump::Delegate*) 	ipc/glue/MessagePump.cpp:235 	cfi
10 	xul.dll 	MessageLoop::RunInternal() 	ipc/chromium/src/base/message_loop.cc:369 	inlined
10 	xul.dll 	MessageLoop::RunHandler() 	ipc/chromium/src/base/message_loop.cc:362 	cfi
11 	xul.dll 	MessageLoop::Run() 	ipc/chromium/src/base/message_loop.cc:344 	inlined
11 	xul.dll 	XRE_InitChildProcess(int, char**, XREChildData const*) 	toolkit/xre/nsEmbedFunctions.cpp:590 	inlined
11 	xul.dll 	mozilla::BootstrapImpl::XRE_InitChildProcess(int, char**, XREChildData const*) 	toolkit/xre/Bootstrap.cpp:60 	cfi
12 	firefox.exe 	NS_internal_main(int, char**, char**) 	browser/app/nsBrowserApp.cpp:397 	inlined
12 	firefox.exe 	wmain(int, wchar_t**) 	toolkit/xre/nsWindowsWMain.cpp:151 	cfi
13 	firefox.exe 	invoke_main() 	/builds/worker/workspace/obj-build/browser/app/D:/a/_work/1/s/src/vctools/crt/vcstartup/src/startup/exe_common.inl:90 	inlined
13 	firefox.exe 	__scrt_common_main_seh() 	/builds/worker/workspace/obj-build/browser/app/D:/a/_work/1/s/src/vctools/crt/vcstartup/src/startup/exe_common.inl:288 	cfi
14 	kernel32.dll 	BaseThreadInitThunk 		cfi
15 	ntdll.dll 	RtlUserThreadStart 		cfi
Duplicate of this bug: 1989702
See Also: → 1996653

I filed bug 1996653 for this signature on MacOS (where GPU process is enabled only in Nightly).

But the big uptick recently is specific to Android, which seems likely to be a separate issue. https://crash-stats.mozilla.org/signature/?signature=IPCError-browser%20%7C%20GPUProcessKill&date=%3E%3D2025-07-27T18%3A08%3A00.000Z&date=%3C2025-10-27T18%3A08%3A00.000Z#graphs (the default graph is by product, the graph by platform is also illustrative).

I had a look through the recent android crashes and couldn't see any obvious reason for a spike - they are on a range of devices, and they follow the usual pattern of the compositor thread waiting on the renderer thread, either for a pause or a readback, etc. And the renderer thread stacks are garbage. And there were no shader compilation annotations which has been a previous reason for a spike.

However, it looks like the recent spike in this signature on android exactly correlates with the decline in reports in bug 1900134. When we first started reporting these timeouts I wasn't sure why the android signatures came under IPCError-content and non-android under IPCError-browser, but for whatever reason it looks like we're still receiving roughly the same number of reports but android ones are now appearing as IPCError-browser as well.

Windows 10, 64 bit,
Firefox 64bit 151.0b3 Crash Report [@ IPCError-browser | GPUProcessKill ]
https://crash-stats.mozilla.org/report/index/e7f6c76e-c828-40d0-8a9e-3fdbc0260529

Crashing Thread (0), Name: MainThread
Frame 	Module 	Signature 	Source 	Trust
0 	win32u.dll 	ZwUserMsgWaitForMultipleObjectsEx 		context
1 	user32.dll 	RealMsgWaitForMultipleObjectsEx(unsigned long, void* const*, unsigned long, unsigned long, unsigned long) 		cfi
2 	xul.dll 	mozilla::widget::WinUtils::WaitForMessage(unsigned long) 	widget/windows/WinUtils.cpp:406 	inlined
2 	xul.dll 	nsAppShell::ProcessNextNativeEvent(bool) 	widget/windows/nsAppShell.cpp:795 	cfi
3 	xul.dll 	nsBaseAppShell::DoProcessNextNativeEvent(bool) 	widget/nsBaseAppShell.cpp:134 	inlined
3 	xul.dll 	nsBaseAppShell::OnProcessNextEvent(nsIThreadInternal*, bool) 	widget/nsBaseAppShell.cpp:270 	inlined
3 	xul.dll 	nsThread::ProcessNextEvent(bool, bool*) 	xpcom/threads/nsThread.cpp:1118 	inlined
3 	xul.dll 	NS_ProcessNextEvent(nsIThread*, bool) 	xpcom/threads/nsThreadUtils.cpp:465 	cfi
4 	xul.dll 	mozilla::ipc::MessagePump::Run(base::MessagePump::Delegate*) 	ipc/glue/MessagePump.cpp:105 	cfi
5 	xul.dll 	MessageLoop::RunInternal() 	ipc/chromium/src/base/message_loop.cc:371 	inlined
5 	xul.dll 	MessageLoop::RunHandler() 	ipc/chromium/src/base/message_loop.cc:364 	cfi
6 	xul.dll 	MessageLoop::Run() 	ipc/chromium/src/base/message_loop.cc:346 	inlined
6 	xul.dll 	nsBaseAppShell::Run() 	widget/nsBaseAppShell.cpp:151 	cfi
7 	xul.dll 	nsAppShell::Run() 	widget/windows/nsAppShell.cpp:672 	cfi
8 	xul.dll 	XRE_RunAppShell() 	toolkit/xre/nsEmbedFunctions.cpp:652 	inlined
8 	xul.dll 	mozilla::ipc::MessagePumpForChildProcess::Run(base::MessagePump::Delegate*) 	ipc/glue/MessagePump.cpp:233 	cfi
9 	xul.dll 	MessageLoop::RunInternal() 	ipc/chromium/src/base/message_loop.cc:371 	inlined
9 	xul.dll 	MessageLoop::RunHandler() 	ipc/chromium/src/base/message_loop.cc:364 	cfi
10 	xul.dll 	MessageLoop::Run() 	ipc/chromium/src/base/message_loop.cc:346 	inlined
10 	xul.dll 	XRE_InitChildProcess(int, char**, XREChildData const*) 	toolkit/xre/nsEmbedFunctions.cpp:590 	inlined
10 	xul.dll 	mozilla::BootstrapImpl::XRE_InitChildProcess(int, char**, XREChildData const*) 	toolkit/xre/Bootstrap.cpp:59 	cfi
11 	firefox.exe 	NS_internal_main(int, char**, char**) 	browser/app/nsBrowserApp.cpp:466 	inlined
11 	firefox.exe 	wmain(int, wchar_t**) 	toolkit/xre/nsWindowsWMain.cpp:150 	cfi
12 	firefox.exe 	invoke_main() 	/builds/worker/workspace/obj-build/browser/app/D:/a/_work/1/s/src/vctools/crt/vcstartup/src/startup/exe_common.inl:90 	inlined
12 	firefox.exe 	__scrt_common_main_seh() 	/builds/worker/workspace/obj-build/browser/app/D:/a/_work/1/s/src/vctools/crt/vcstartup/src/startup/exe_common.inl:288 	cfi
13 	kernel32.dll 	BaseThreadInitThunk 		cfi
14 	ntdll.dll 	RtlUserThreadStart 		cfi

=================

And some Seconds later with same stack trace as far as I can see..:

https://crash-stats.mozilla.org/report/index/fb5f3d61-59ef-4f0f-a99a-907260260529

Based on the topcrash criteria, the crash signature linked to this bug is not a topcrash signature anymore.

For more information, please visit BugBot documentation.

since I cannot edit the see also information, here are some additions:

  • bug 2072450 (Crash Report [@ IPCError-browser | GPUProcessKill ] Thunderbird, Windows 10)
  • bug 2072864 (Crash Report [@ IPCError-browser | GPUProcessKill ] Firefox, Windows 10)
Duplicate of this bug: 2072864

Crash report: https://crash-stats.mozilla.org/report/index/1a63c803-e47f-4ae6-90dc-ad4370261006

Crash Reason:

EXC_BREAKPOINT / EXC_ARM_BREAKPOINT at 0x0000000184607c34

Top 10 frames of the hung main thread (nothing crashed here — a watchdog killed the process; these are what the main thread is waiting on):

0  libsystem_kernel.dylib  mach_msg2_trap
1  libsystem_kernel.dylib  mach_msg2_internal
2  libsystem_kernel.dylib  mach_msg_overwrite
3  libsystem_kernel.dylib  mach_msg
4  CoreFoundation  __CFRunLoopServiceMachPort
5  CoreFoundation  __CFRunLoopRun
6  CoreFoundation  _CFRunLoopRunSpecificWithOptions
7  HIToolbox  RunCurrentEventLoopInMode
8  HIToolbox  ReceiveNextEventCommon
9  HIToolbox  _BlockUntilNextEventMatchingListInMode

Filed because this signature's crash volume spiked. 76 distinct installations hit this signature in the last 7 days (83 reports), against 17.27 expected from its own rate over the preceding 56 days -- 4.4x, normalised for the channel's daily installation count. This signature is not new: its first report is in build 20240514215634 (2024-05-14), 874 days before this build.

Clouseau analysis (automated -- nothing below was checked by a human; a claim it could not ground is deliberately absent):

These reports are written when the parent process kills a hung GPU process: a synchronous IPC to the compositor gets no reply within layers.gpu-process.ipc_reply_timeout_ms (10 s), the extension of up to 20 s runs out, and CompositorManagerChild::ShouldContinueFromReplyTimeout calls GPUProcessManager::KillProcess(true). On nightly the 'spike' looks like reporting coming back, not more hangs. Builds from about 2026-06-22 to 2026-09-17 sent almost no reports of this signature, while June builds sent about 15-25 a day. Reports came back at a similar level starting with the first build after bug 1949178 (40a64974956f, 2026-09-17), which moved CrashReporter::CreateMinidumpsAndPair (the call that captures the killed GPU process's dump) to the out-of-process crash helper. The suggested action is to treat the nightly rise as a collection artefact: crash-reporting owners should confirm that kill-time paired minidumps were being lost from about build 20260622 on (bug 1955963 is the likely cause), and the GPU hang itself should be followed in the existing graphics bugs, not attributed to anything in the 2026-10-06 pushlog.

Starting point -- a candidate, NOT an established cause (medium confidence): 40a64974956f (bug 1949178) by :gsvelto.

The diff replaces the old in-process paired-dump path in CrashReporterHost::GenerateMinidumpAndPair / CreateMinidumpsAndPair (Breakpad WriteMinidumpForChild plus AddSharedAnnotations on mExtraAnnotations) with generate_crash_report(gCrashHelperClient, aId, ...). Annotations are now read from the helper-written .extra file (ReadExtraFile) and AddCommonAnnotations is applied. Bug 1955963 (77310ded9dfc, 2026-06-22) earlier changed AddSharedAnnotations so it no longer copies main-process annotations into the target table. The reporting gap starts right after that (last reporting build 20260621215606) and ends with the first build after 40a64974956f (20260918092903). The new report key set (CrashEventID, OS, OSVersion, CPUArchitecture, GpuSandboxLevel at 100% after the split) fits the new path. This change restored reporting; it did not create GPU hangs. I did not show directly that reports from the gap were rejected or dropped.

Possible path to the crash:

Mechanism (observed in source at pinned rev 953be279951f): parent-side CompositorManagerChild::ShouldContinueFromReplyTimeout logs 'Killing GPU process due to IPC reply timeout' and calls GPUProcessManager::KillProcess(/* aGenerateMinidump */ true) once the sync IPC has waited past the 10000 ms timeout plus up to 20000 ms of extension. Many reports carry this exact graphics_critical_error note. The GPU process main thread is idle in its event loop because the hang is on another thread. Path for the representative macOS report (from the brief and the earlier pipeline run, not re-read here): the Compositor is in SynchronousTask::Wait, waiting on the Renderer, which is inside the system call IOSurfaceCreate. The Windows cohort (about 80% of nightly reports after the split) was not inspected, so its blocking thread is unknown.

Population (derived): reports stopped from builds after about 20260621 and resumed from 20260918092903 at a level similar to June. Inferred (medium): the gap was lost reports, not missing hangs. It lines up with the crash-reporting changes in bug 1955963 (2026-06-22) and bug 1949178 (2026-09-17). Trigger for the hangs themselves: unknown.

Checked:

  • On nightly, versions 155-157 have essentially no reports of this signature over 120 days (one 155.0a1 report on 2026-07-28). 153/154 had about 15-29 per day in June, and 158/159 run about 5-21 per day from 2026-09-18. (facets version by_day, nightly, 120d)
  • Before the 20260918 split, the top builds are all 20260605-20260621 or older; after it, the earliest is 20260918092903 (8 reports). (facets build_id split_at_build=20260918000000, nightly, 120d)
  • Builds from 20260918 have CrashEventID, CPUArchitecture, OS, OSVersion and GpuSandboxLevel in 100% of crash_report_keys; none of these is in the before-split top-20 list. (facets crash_report_keys split_at_build=20260918000000, nightly, 60d)
  • 40a64974956f (bug 1949178, 2026-09-17) rewrites CreateMinidumpsAndPair to generate the child dump through gCrashHelperClient and read its annotations from the .extra file, replacing Breakpad WriteMinidumpForChild plus AddSharedAnnotations. (mcp__patch__diff 40a64974956f: toolkit/crashreporter/nsExceptionHandler.cpp, ipc/glue/CrashReporterHost.cpp/.h)
  • 77310ded9dfc (bug 1955963, 2026-06-22) removes the code in AddSharedAnnotations that filled in main-process annotations missing from the passed-in table, leaving only filtering. (mcp__patch__diff 77310ded9dfc: toolkit/crashreporter/nsExceptionHandler.cpp hunk @@ -3109)
  • Bug 1949178 ('Re-enable capturing minidumps of child shutdown hangs after OOP crash generation lands') has a known regression, bug 2078233 ('Increase in missing additional minidump errors when submitting crashes'). (Bugzilla 1949178, 2078233)
  • The GPU kill happens in ShouldContinueFromReplyTimeout after the layers.gpu-process.ipc_reply_timeout_ms (10000) timeout plus up to extend_ipc_reply_timeout_ms (20000) of extension, and calls KillProcess(true). (gfx/layers/ipc/CompositorManagerChild.cpp@953be279951f:229-262; modules/libpref/init/StaticPrefList.yaml@953be279951f:10510-10518)
  • CompositorManagerChild.cpp has had no functional change since 2026-04 (only include sorting). (file_history gfx/layers/ipc/CompositorManagerChild.cpp)

Alternatives considered:

  • Changesets in the 2026-10-06 pushlog window (including 7d3750ea4c76 / 5e9eb3e8550d, bug 1564615): disfavored -- the rise begins with build 20260918092903, weeks before that window, and 159 is within 2x of 158 by per-version share.
  • A change to the GPU kill/timeout logic: not supported -- CompositorManagerChild.cpp is unchanged except for include sorting since April, and the timeout prefs read 10000/20000 at the build rev (blame of the pref file timed out, so history of the values is unverified).
  • One GPU vendor or driver as the trigger: disfavored -- after the split, reports spread across NVIDIA, AMD, Intel and Apple with many driver versions.
  • Hardware error: disfavored -- 0% bit-flip and 2% Raptor Lake share per the brief.
  • Release-channel burst on 2026-09-27..10-01 (mostly 'Unknown' OS) as the same phenomenon: unresolved -- not investigated; it is a different channel and population.

Worth checking first:

  • Were GPUProcessKill paired reports from nightly builds 20260622-20260917 written but rejected by Socorro, or never written? Crash-reporting owners can check the old CreateMinidumpsAndPair + AddSharedAnnotations path after 77310ded9dfc for missing ProductName/BuildID/ReleaseChannel, or check crash-ping (Glean) counts of GPU-process kills over that period.
  • What the GPU process is blocked on in the Windows cohort (~80% of nightly reports): read the Compositor and Renderer threads of several Windows 11 reports from builds >= 20260918.
  • Whether the true GPU-hang rate changed between June and now: compare telemetry for GPU-process kills or IPC reply timeouts, which does not depend on minidump submission.
  • Whether bug 2078233 (missing additional minidump after 1949178) affects these paired reports: check DumperError / additional-minidump presence on this signature.

:gsvelto, can you have a look please?

Filed automatically by Clouseau, which analyses nightly crashes with an LLM. Nothing above was written or checked by a human. Please close it as INVALID if it is wrong — that is useful feedback, not a nuisance.

This is not an independent Mozilla discovery: the crash is known here only because somebody submitted the crash report linked at the top. If this duplicates an existing report, please resolve THIS bug as the duplicate and leave the credit with the earlier reporter.

Flags: needinfo?(gsvelto)

Bug 1955963 caused these reports to not be generated until bug 1949178 landed, so that's the most likely explanation.

Flags: needinfo?(gsvelto)
You need to log in before you can comment on or make changes to this bug.