Closed Bug 49108 Opened 25 years ago Closed 25 years ago

N601 (Linux) crash in [@ libc.so.6 - nsDiskCacheRecord::RetrieveInfo]

Categories

(Core :: Networking: Cache, defect, P3)

x86
Linux
defect

Tracking

()

VERIFIED FIXED

People

(Reporter: jwbaker, Assigned: neeti)

References

()

Details

(Keywords: crash, topcrash, Whiteboard: [dogfood-][nsbeta3+])

Crash Data

Attachments

(1 file)

I'm crashing all over the place wherever I browse to today. I got the same stack at http://www.compaq.com and http://www.arstechnica.com/. The stack points to disk cache code. I got this stack trace on a Linux debug build pulled 2000-08-15-09-ish: #0 0x403a2d0f in memcpy () from /lib/libc.so.6 #1 0x40df4c03 in nsDiskCacheRecord::RetrieveInfo (this=0x893b118, aInfo=0x8748b2b, aInfoLength=463) at nsDiskCacheRecord.cpp:447 #2 0x40df1575 in nsDBEnumerator::GetNext (this=0x872a690, _retval=0xbffff268) at nsDBEnumerator.cpp:100 #3 0x40dee426 in nsReplacementPolicy::AddAllRecordsInCache (this=0x86b3990, aCache=0x85987a0) at nsReplacementPolicy.cpp:170 #4 0x40def2d8 in nsReplacementPolicy::LoadAllRecordsInAllCacheDatabases (this=0x86b3990) at nsReplacementPolicy.cpp:619 #5 0x40def34a in nsReplacementPolicy::Evict (this=0x86b3990, aTargetOccupancy=4849) at nsReplacementPolicy.cpp:646 #6 0x40de926b in nsCacheManager::LimitDiskCacheSize (skipCheck=1) at nsCacheManager.cpp:524 #7 0x40ded7b1 in InterceptStreamListener::OnStopRequest (this=0x87e68b8, channel=0x88ba710, ctxt=0x0, aStatus=0, aStatusArg=0x401698e4) at nsCachedNetData.cpp:1178 #8 0x40e13c1a in nsHTTPChannel::ResponseCompleted (this=0x88ba710, aListener=0x87e68b8, aStatus=0, aStatusArg=0x401698e4) at nsHTTPChannel.cpp:1761 #9 0x40e1e4d8 in nsHTTPServerListener::OnStopRequest (this=0x8859bd8, channel=0x8875a6c, i_pContext=0x88ba710, i_Status=0, aStatusArg=0x401698e4) at nsHTTPResponseListener.cpp:719 #10 0x40db5d76 in nsOnStopRequestEvent::HandleEvent (this=0x87c1498) at nsAsyncStreamListener.cpp:301 #11 0x40db5317 in nsStreamListenerEvent::HandlePLEvent (aEvent=0x86c2f28) at nsAsyncStreamListener.cpp:97 #12 0x4011730f in PL_HandleEvent (self=0x86c2f28) at plevent.c:587 #13 0x401171b1 in PL_ProcessPendingEvents (self=0x80d3078) at plevent.c:528 #14 0x40118f31 in nsEventQueueImpl::ProcessPendingEvents (this=0x80d3050) at nsEventQueue.cpp:356 #15 0x409deec8 in event_processor_callback (data=0x80d3050, source=8, condition=GDK_INPUT_READ) at nsAppShell.cpp:158 #16 0x409deb07 in our_gdk_io_invoke (source=0x81d4f88, condition=G_IO_IN, data=0x81d4f78) at nsAppShell.cpp:58 #17 0x40b9920e in g_io_unix_dispatch (source_data=0x81d4fa0, current_time=0xbffff65c, user_data=0x81d4f78) at giounix.c:135 #18 0x40b9a717 in g_main_dispatch (dispatch_time=0xbffff65c) at gmain.c:656 #19 0x40b9acdb in g_main_iterate (block=1, dispatch=1) at gmain.c:877 #20 0x40b9ae59 in g_main_run (loop=0x81d4fe8) at gmain.c:935 #21 0x40acc069 in gtk_main () at gtkmain.c:476 #22 0x409df5b1 in nsAppShell::Run (this=0x810afe8) at nsAppShell.cpp:335 #23 0x40510388 in nsAppShellService::Run (this=0x810f738) at nsAppShellService.cpp:378 #24 0x805558c in main1 (argc=1, argv=0xbffff964, nativeApp=0x0) at nsAppRunner.cpp:943 #25 0x8055c70 in main (argc=1, argv=0xbffff964) at nsAppRunner.cpp:1123 #26 0x4035e2e7 in __libc_start_main () from /lib/libc.so.6
Keywords: crash
This may be a regression of bug 40084, which has a similar stack trace.
I am seeing this as well on my Linux optimized build from tree closure 8-17. I think it happens when the cache reaches a certain state (perhaps when it's full???). Nominating for dogfood. Once the browser reaches this state I can't do anything useful. I have saved my ~/.mozilla/ directory from the state when the crash started so that you can use it for debugging if you want. I can email it to you (I don't want it public for the whole world, though).
Keywords: dogfood
Seeing this in my debug build from the same time: #0 0x403b44a7 in memcpy (dstpp=0x42215008, srcpp=0x8881dc8, len=6357060) at ../sysdeps/generic/memcpy.c:55 dstpp = (void *) 0x42215008 len = 6357060 dstp = 1109578304 srcp = 143237120 #1 0x40e46131 in nsDiskCacheRecord::RetrieveInfo (this=0x8868820, aInfo=0x887c983, aInfoLength=282) at /home/david/mozilla/src/mozilla/netwerk/cache/filecache/nsDiskCacheRecord.cpp:459 cur_ptr = 0x8881dc8 "v" file_url = 0x42215008 "v" name_len = 6357060 id = 21299226 #2 0x40e42a11 in nsDBEnumerator::GetNext (this=0x8861878, _retval=0xbffff220) at /home/david/mozilla/src/mozilla/netwerk/cache/filecache/nsDBEnumerator.cpp:100 this = (nsDBEnumerator *) 0x8861878 rv = 1089417700 #3 0x40e3f8a2 in nsReplacementPolicy::AddAllRecordsInCache (this=0x8816a70, aCache=0x86b0918) at /home/david/mozilla/src/mozilla/netwerk/cache/mgr/nsReplacementPolicy.cpp:170 notDone = 1 rv = 0 iterator = {mRawPtr = 0x8861878} iSupports = {<nsCOMPtr_base> = {mRawPtr = 0x0}, <No data fields>} record = {mRawPtr = 0x8868820} #4 0x40e4075c in nsReplacementPolicy::LoadAllRecordsInAllCacheDatabases ( this=0x8816a70) at /home/david/mozilla/src/mozilla/netwerk/cache/mgr/nsReplacementPolicy.cpp:619 this = (nsReplacementPolicy *) 0x8816a70 rv = 135585368 cacheInfo = (CacheInfo *) 0x87a8e58 #5 0x40e407ca in nsReplacementPolicy::Evict (this=0x8816a70, aTargetOccupancy=4849) at /home/david/mozilla/src/mozilla/netwerk/cache/mgr/nsReplacementPolicy.cpp:646 this = (nsReplacementPolicy *) 0x8816a70 i = 135585368 entry = (nsCachedNetData *) 0x40e407a0 rv = 3199 occupancy = 0 truncatedLength = 1073785200 record = {mRawPtr = 0x0} #6 0x40e3a657 in nsCacheManager::LimitDiskCacheSize (skipCheck=1) at /home/david/mozilla/src/mozilla/netwerk/cache/mgr/nsCacheManager.cpp:524 rv = 0 spaceManager = (nsReplacementPolicy *) 0x8816a70 occupancy = 4901 diskCacheCapacity = 5000 #7 0x40e3ec15 in InterceptStreamListener::OnStopRequest (this=0x8861918, channel=0x8863828, ctxt=0x0, aStatus=0, aStatusArg=0x4016e8e4) at /home/david/mozilla/src/mozilla/netwerk/cache/mgr/nsCachedNetData.cpp:1193 channel = (nsIChannel *) 0x8863828
Incidentally, the "v" that cur_ptr points to is pointing to the middle of the contents of one of my web pages (as a 2-byte string): vid Baron">send me e-mail</a>. It is not possible for This is part of the page ( http://www.people.fas.harvard.edu/~dbaron/ ) that I was trying to load when I took this stack trace. (The stuff before the v is also good web page data.) So, is cur_ptr pointing to the middle of something else?
Doing more poking around in ::RetrieveInfo: (gdb) p ((char*)aInfo)+2 $91 = 0x887c985 "/home/david/.mozilla/David-1/Cache/0a/012c002a" (gdb) p ((char*)aInfo)+65 $118 = 0x887c9c4 "http://geography.uoregon.edu/envchange/clim_animations/thumbnails/lhtfl_clim01sm.gifZ\001" (gdb) p ((char*)aInfo)+154 $127 = 0x887ca1d "HTTP headers" (gdb) p ((char*)aInfo)+170 $133 = 0x887ca2d "(HTTP/1.1 200 OK\r\nserver: Microsoft-IIS/4.0\r\ncache-control: max-age=1800\r\nexpires: Tue, 15 Aug 2000 17:50:55 GMT\r\ndate: Tue, 15 Aug 2000 17:20:55 GMT\r\ncontent-type: image/gif\r\naccept-ranges: bytes\r\nla"... (gdb) p cur_ptr - aInfo $136 = 21573 (gdb) p mKeyLength $137 = 48 (gdb) p mKey $141 = 0x884d328 "http://netscape.weather.com/includes/header.html8" (gdb) p mMetaDataLength $138 = 21505 (gdb) p mMetaData $139 = 0x88944f8 "" (gdb) p aInfoLength $142 = 282 It looks like mMetaDataLength is wrong. Furthermore, I wonder why the stuff I see in aInfo is different from what I see in mKey.
Keywords: nsbeta3
If we ever fail on one of these sanity checks, we should probably assume the whole cache is corrupt, toss it, and re-create it.
Putting on [dogfood-] radar.
Whiteboard: [dogfood-]
this is very common and very bad. must at least have a hack fix that tosses away corrupt cache. plus.
Whiteboard: [dogfood-] → [dogfood-][nsbeta3+]
*** Bug 50077 has been marked as a duplicate of this bug. ***
Marking topcrash. This is one of the top crashers in Windows talkback reports of the past week.
Keywords: topcrash
neeti, could you review the paranoia code? that will at least stop us from crashing.
Reviewed the patch. Looks good. r=neeti.
*** Bug 50428 has been marked as a duplicate of this bug. ***
Checked in Waterson's patch.
Status: NEW → RESOLVED
Closed: 25 years ago
Resolution: --- → FIXED
Am I correct that this patch means that once the cache is corrupted, it just stops being used and silently disappears? That seems to be a serious bug too (although not as bad as crashing). Is one filed?
We do not have a specific bug to throw away the cache once it is corrupted. Bug 47403 will help us solve cache corruption problems. Neeti
Filed bug 50559 regarding what to do with a corrupt cache.
verified: Linux rh6 2000091408
Status: RESOLVED → VERIFIED
updated summary with N601, this is a topcrasher with the N601 release under the stack signature libc.so.6. adding [@ libc.so.6 - nsDiskCacheRecord::RetrieveInfo] for tracking. leaving verified fixed for now, since it has been fixed on trunk.
Summary: Crash in nsDiskCacheRecord::RetrieveInfo → N601 (Linux) crash in [@ libc.so.6 - nsDiskCacheRecord::RetrieveInfo]
Crash Signature: [@ libc.so.6 - nsDiskCacheRecord::RetrieveInfo]
You need to log in before you can comment on or make changes to this bug.

Attachment

General

Creator:
Created:
Updated:
Size: