Closed Bug 951846 Opened 12 years ago Closed 5 years ago

[FlatFish] ClientThebesLayer::PaintThebes() costs a lot on RotatedBuffer::DrawBufferWithRotation

Categories

(Core :: Graphics, defect)

ARM
Gonk (Firefox OS)
defect
Not set
normal

Tracking

()

RESOLVED WONTFIX

People

(Reporter: vlin, Unassigned)

References

Details

Attachments

(1 file)

It takes ~55ms to do RotatedBuffer::DrawBufferWithRotation() which is several times of other HD display device.
Blocks: 937713
No longer blocks: 937713
Blocks: 937713
It turns out the slowness is due to the performance of memcpy between two GraphicBuffer is so poor. Our experiment shows that it takes more than 90ms to memcpy 1280x800 buffer from A to B GraphicBuffer. Need feedback from partner's vendor.
Flags: needinfo?(phterry)
usage = GraphicBuffer::USAGE_SW_READ_OFTEN | GraphicBuffer::USAGE_SW_WRITE_OFTEN |GraphicBuffer::USAGE_HW_TEXTURE;
Flags: needinfo?(phterry)
(In reply to Vincent Lin[:vilin] from comment #1) > It turns out the slowness is due to the performance of memcpy between two > GraphicBuffer is so poor. Our experiment shows that it takes more than 90ms > to memcpy 1280x800 buffer from A to B GraphicBuffer. Need feedback from > partner's vendor. Is there any testing patch or trace data compared to other platform that we can refer to?
Flags: needinfo?(vlin)
Please check the test code for GraphicBuffer in GonkDisplayJB.cpp/h. Each memcpy takes 90ms. BTW. There is also a test code in RotatedBuffer.cpp for locally-created buffer. 1st memcpy takes 18ms while 2nd memcpy takes 1~2ms. memcpy GraphicBuffer performance affects DrawBufferQuadrant's cost a lot which is frequently called in rendering flow. We noticed that for HD buffer flatfish costs 50~60ms while N4 costs 5ms only.
Flags: needinfo?(vlin)
That's quite interesting for the locally-created buffer tests. Here are my thoughts (just guesses). I think it's not about cache-miss issue, since the buffer size is quite large. Is it possible that the longer time (18ms) taken by the 1st memcpy() was due to the virtual memory management and the copy-on-write mechanism? When the pages happen to be allocated for the very first time in the heap, the brk() system calls will be invoked to increase the heap size. Due to security reasons, the kernel will allocate the newly created pages as clean pages. It simply sets the page entries to make it point the zero page (read-only page 0). After that, when these pages (specifically, pages of buffer2) are to be written as the destination of the memcpy() call, page faults occur due to the read-only attribute of the zero page. The page fault handler, each time called, will allocate a physical page from the free pages and initialize it with zeros. Then, the memcpy() continues its execution and fill (write) the page with the content of a source page (from buffer1, which may actually contained just zeros in this situation.) This process repeats until all pages of buffer2 get copied from buffer1. Due to this copy-on-write behavior, the 1st memcpy() takes much longer time. The 2nd memcpy() doesn't have that problem because the physically pages of buffer2 are already there. FYR.
(In reply to William Liang from comment #5) > That's quite interesting for the locally-created buffer tests. Here are my > FYR. Thanks for the sharing. Locally-created buffer test is just for comparison with GraphicBuffer test. We concerned more about the fact memcpy() on GraphicBuffer is much longer because Gecko utilizes Cairo to copy GraphicBuffer a lot.
Terry, Would you please help to verify it on your Android device with our test code. Supposedly, you will get the same result. Also, I'm curious about the skia performance on Flatfish running Android. Or there is no chance for skia to paint/swap a full frame in Android ? So it never encountered this worse case. Verifying the feedback from IMG, but there's still no improvement. (Only set READ_OFTEN flag to src buffer, WRITE_OFTEN flag to dst buffer.) The detailed feedback is as follows. For the test you are running I'd recommend to set the usage bits as READ_OFTEN for the source buffer, and only WRITE_OFTEN for the destination. You don't need TEXTURE for either lock. Otherwise our driver may flush the CPU cache for the source buffer, which is pointless. Having said that, due to the non-coherent GPU/CPU cache there is an extra cost associated with allocations (we have to allocate GPU pages and do the CPU mappings for the newly allocated memory).
Flags: needinfo?(phterry)
Component: General → Graphics
Product: Firefox OS → Core
Status: NEW → RESOLVED
Closed: 5 years ago
Flags: needinfo?(phterry)
Resolution: --- → WONTFIX
You need to log in before you can comment on or make changes to this bug.

Attachment

General

Created:
Updated:
Size: