Open Bug 1954570 Opened 1 year ago Updated 1 year ago

Trigger GC based on utilization of GPU resources via WebGPU

Categories

(Core :: Graphics: WebGPU, task, P3)

task

Tracking

()

People

(Reporter: aleiserson, Unassigned)

References

(Blocks 1 open bug)

Details

GPU resources allocated by WebGPU are kept alive based on the state of associated Javascript objects. GPU memory pressure will not necessarily correlate with main memory pressure or other existing GC triggers. (I am not sure exactly what GPU state is wired to GC currently, but regardless, those GPU contexts are separate from WebGPU's GPU contexts, so will not cover WebGPU utilization.)

This is most apparent today with issues like https://bugzilla.mozilla.org/show_bug.cgi?id=1914934 where there is runaway GPU resource consumption, but it may be possible to see issues here in any application that allocates substantial GPU resources on an ongoing basis (e.g. per frame), when there is not much other activity that will trigger GC.

https://fluid.loga.nz/, bug 1927126, is another test case for this.

I asked :sfink for advice on this. He says it seems like we should use AddAssociatedMemory.

Here's an an example of how it was done for canvas.

He says we might want to add our own memory type.

AddAssociatedMemory seems to only be intended for CPU memory.

We'd probably need to add a new set of size and threshold for GPU memory here: https://searchfox.org/mozilla-central/rev/5fb48bf50516ed2529d533e5dfe49b4752efb8b8/js/src/gc/ZoneAllocator.h#175-197

but it may be possible to see issues here in any application that allocates substantial GPU resources on an ongoing basis (e.g. per frame), when there is not much other activity that will trigger GC.

While this can happen, it's discouraged in the graphics space.

The WebGPU spec already has a mechanism that allows UAs to "lose the device" in high memory pressure situations; which will effectively sever the connection to the GPU and destroy all resources.

I think we should first implement detection of high memory pressure in wgpu-core and lose the device if a threshold is reached. We can then see if it's still worth communicating GPU memory usage to the GC.

:teoxoy:

The WebGPU spec already has a mechanism that allows UAs to "lose the device" in high memory pressure situations; which will effectively sever the connection to the GPU and destroy all resources.

I don't see this in the spec. Where is that? 👀

Flags: needinfo?(ttanasoaia)

It's not explicitly stated that it's allowed in high memory pressure situations but more broadly:

Any time the user agent needs to revoke access to a device, it calls lose the device(device, "unknown") on the device’s device timeline, potentially ahead of other operations currently queued on that timeline.

from https://www.w3.org/TR/webgpu/#devices

It's use for high memory pressure situations was talked about in some issues/meetings.

Flags: needinfo?(ttanasoaia)

I opened https://github.com/gfx-rs/wgpu/issues/7460 to track the wgpu-core implementation of OOM detection & mitigation.

There are a few ideas that we can pursue here:

  • First, let's make sure that the objects that we want to GC are nursery-allocatable. This will naturally encourage the JS GC to notice that they're unused quickly. We should be able to verify this by simply writing code that creates a lot of the objects and then looking at GC behavior in the Firefox profiler.

  • We can simply add a new MemoryUse for GC memory. As long as it's measured in bytes, it doesn't matter that the number here isn't counting CPU bytes, because these counters are only used to assess the pace of allocation and to compare with other high water marks; they're not expected to be drawn from any particular CPU memory pool.

  • There is an existing MEM_PRESSURE GCReason, but we should be cautious about it, as DoLowMemoryGC may be for emergency situations.

  • If MEM_PRESSURE isn't appropriate, we can consider adding a new GCReason. Check with Steve Fink or Jon Coppeard about this.

You need to log in before you can comment on or make changes to this bug.