Closed Bug 40084 Opened 26 years ago Closed 25 years ago

[CRASH] Crash in disk cache code

Categories

(Core :: Networking: Cache, defect, P1)

All
Linux
defect

Tracking

()

VERIFIED DUPLICATE of bug 50977

People

(Reporter: jce2, Assigned: neeti)

References

()

Details

(Keywords: crash, topcrash, Whiteboard: [nsbeta2+][dogfood-] ETA 7/28)

Core was generated by `./mozilla-bin'. Program terminated with signal 6, Aborted. (gdb) backtrace backtrace #0 0x405e551d in __mmap () from /lib/libc.so.6 #1 0x40595e4d in chunk_alloc (ar_ptr=0x40620a00, nb=1701406320) at malloc.c:1809 #2 0x40595972 in __libc_malloc (bytes=1701406313) at malloc.c:2596 #3 0x4026aacc in PR_Malloc (size=1701406313) at prmem.c:38 #4 0x4011baac in nsAllocatorImpl::Alloc (this=0x80663e8, size=1701406313) at nsAllocator.cpp:82 #5 0x4011bb7a in nsAllocator::Alloc (size=1701406313) at nsAllocator.cpp:148 #6 0x40dee049 in nsDiskCacheRecord::RetrieveInfo (this=0x95a8528, aInfo=0x8742fb0, aInfoLength=386) at nsDiskCacheRecord.cpp:442 #7 0x40deacf5 in nsDBEnumerator::GetNext (this=0x95a7ba8, _retval=0xbffff260) at nsDBEnumerator.cpp:100 #8 0x40de7a8a in nsReplacementPolicy::AddAllRecordsInCache (this=0x866cb98, aCache=0x8609128) at nsReplacementPolicy.cpp:170 #9 0x40de8974 in nsReplacementPolicy::LoadAllRecordsInAllCacheDatabases (this=0x866cb98) at nsReplacementPolicy.cpp:619 #10 0x40de89e9 in nsReplacementPolicy::Evict (this=0x866cb98, aTargetOccupancy=4849) at nsReplacementPolicy.cpp:646 #11 0x40de2969 in nsCacheManager::LimitDiskCacheSize () at nsCacheManager.cpp:516 #12 0x40de299b in nsCacheManager::LimitCacheSize () at nsCacheManager.cpp:526 #13 0x40de98dd in CacheOutputStream::Write (this=0x8b026c0, aBuf=0x9375800 "-article.cgi?article_key=1000351120\">2. oh......</A>(aquiri, a.k.a. extreme burger, no happy meals here / 04.19.00)<DD>\n<DL>\n<DT><A HREF=\"read-article.cgi?article_key=1000351176\">1. You aren't doing y"..., aCount=1460, aActualBytes=0xbffff374) at nsCacheEntryChannel.cpp:89 #14 0x40de6b34 in InterceptStreamListener::write (this=0x918fa38, aBuf=0x9375800 "-article.cgi?article_key=1000351120\">2. oh......</A>(aquiri, a.k.a. extreme burger, no happy meals here / 04.19.00)<DD>\n<DL>\n<DT><A HREF=\"read-article.cgi?article_key=1000351176\">1. You aren't doing y"..., aNumBytes=1460) at nsCachedNetData.cpp:1192 #15 0x40de6bc8 in InterceptStreamListener::Read (this=0x918fa38, buf=0x9375800 "-article.cgi?article_key=1000351120\">2. oh......</A>(aquiri, a.k.a. extreme burger, no happy meals here / 04.19.00)<DD>\n<DL>\n<DT><A HREF=\"read-article.cgi?article_key=1000351176\">1. You aren't doing y"..., count=1460, aActualBytes=0xbffff490) at nsCachedNetData.cpp:1180 #16 0x417bd796 in nsParser::OnDataAvailable (this=0x93382e0, channel=0x8e153b8, aContext=0x0, pIStream=0x918fa3c, sourceOffset=0, aLength=1460) at nsParser.cpp:1840 #17 0x40ee7a70 in nsDocumentOpenInfo::OnDataAvailable (this=0x8e15130, aChannel=0x8e153b8, aCtxt=0x0, inStr=0x918fa3c, sourceOffset=0, count=1460) at nsURILoader.cpp:189 #18 0x40e16350 in nsHTTPFinalListener::OnDataAvailable (this=0x8e151a8, aChannel=0x8e153b8, aContext=0x0, aStream=0x918fa3c, aSourceOffset=0, aCount=1460) at nsHTTPResponseListener.cpp:1216 #19 0x40de6d2f in InterceptStreamListener::OnDataAvailable (this=0x918fa38, channel=0x8e153b8, ctxt=0x0, inStr=0x91ddd5c, sourceOffset=0, count=1460) at nsCachedNetData.cpp:1164 #20 0x40e1486d in nsHTTPServerListener::OnDataAvailable (this=0x9280700, channel=0x8de5964, context=0x8e153b8, i_pStream=0x91ddd5c, i_SourceOffset=471892, i_Length=1460) at nsHTTPResponseListener.cpp:540 #21 0x40db0e34 in nsOnDataAvailableEvent::HandleEvent (this=0x95a83d8) at nsAsyncStreamListener.cpp:406 #22 0x40db0056 in nsStreamListenerEvent::HandlePLEvent (aEvent=0x95a8400) at nsAsyncStreamListener.cpp:97 #23 0x40110e93 in PL_HandleEvent (self=0x95a8400) at plevent.c:575 #24 0x40110d2d in PL_ProcessPendingEvents (self=0x811bd28) at plevent.c:520 #25 0x40112a91 in nsEventQueueImpl::ProcessPendingEvents (this=0x8129330) at nsEventQueue.cpp:316 #26 0x40a6344b in event_processor_callback (data=0x8129330, source=6, condition=GDK_INPUT_READ) at nsAppShell.cpp:143 #27 0x40a6306a in our_gdk_io_invoke (source=0x81b47d8, condition=G_IO_IN, data=0x81cd040) at nsAppShell.cpp:56 #28 0x40c1d0c7 in g_io_unix_dispatch (source_data=0x81b4838, current_time=0xbffff884, user_data=0x81cd040) at giounix.c:135 #29 0x40c1e6c7 in g_main_dispatch (current_time=0xbffff884) at gmain.c:652 #30 0x40c1ec7b in g_main_iterate (block=1, dispatch=1) at gmain.c:870 #31 0x40c1ee09 in g_main_run (loop=0x81c0308) at gmain.c:928 #32 0x40b4aa47 in gtk_main () at gtkmain.c:475 #33 0x40a63a9a in nsAppShell::Run (this=0x811b9d8) at nsAppShell.cpp:313 #34 0x40700edf in nsAppShellService::Run (this=0x8122608) at nsAppShellService.cpp:371 #35 0x805319b in main1 (argc=1, argv=0xbffffb54, nativeApp=0x0) at nsAppRunner.cpp:904 #36 0x80538c0 in main (argc=1, argv=0xbffffb54) at nsAppRunner.cpp:1188 #37 0x40556ab2 in __libc_start_main (main=0x80536f4 <main>, argc=1, argv=0xbffffb54, init=0x804eda8 <_init>, fini=0x805cf88 <_fini>, rtld_fini=0x4000b0b0 <_dl_fini>, stack_end=0xbffffb4c) at ../sysdeps/generic/libc-start.c:78
Keywords: crash
jce2@po.cwru.edu, could you please explain how to reproduce the bug?
Also what version of the product are you running?
Recommend resolving as WORKSFORME, due to lack of detail.
This is a SPORATIC bug..it's somewhat difficult to reproduce. Try loading long pages and clicking on one of the links while it is still downloading. Many times, the browser will just freeze up and then crash.
jce2@po.cwru.edu, can you give a url of a long page where I can reproduce the crash? Neeti
I have been able to track down some of this Here is some debug output: First one that looks ok: nsDiskCacheRecord::RetrieveInfo(7f20b8, 1a6) mInfoSize=1a6 mRecordID=10a0082 mKeyLength=31 nsDiskCacheRecord::RetrieveInfo1 -> Alloc(31 * 1) mMetaDataLength=13b nsDiskCacheRecord::RetrieveInfo2 -> Alloc(13b * 1) name_len=26 nsDiskCacheRecord::RetrieveInfo3 -> Alloc(26 * 1) nsDiskCacheRecord::GetMetaData -> Alloc(13b * 1) Here are a few not-so-good ones.. nsDiskCacheRecord::RetrieveInfo(7f2c68, 195) mInfoSize=195 mRecordID=1070004 mKeyLength=44 nsDiskCacheRecord::RetrieveInfo1 -> Alloc(44 * 1) mMetaDataLength=31303a31 nsDiskCacheRecord::RetrieveInfo2 -> Alloc(31303a31 * 1) name_len=0 nsDiskCacheRecord::RetrieveInfo3 -> Alloc(0 * 1) nsDiskCacheRecord::RetrieveInfo(7f1d69, 132) mInfoSize=132 mRecordID=1060046 mKeyLength=26 nsDiskCacheRecord::RetrieveInfo1 -> Alloc(26 * 1) mMetaDataLength=21094039 nsDiskCacheRecord::RetrieveInfo2 -> Alloc(21094039 * 1) name_len=dedebdd6 nsDiskCacheRecord::RetrieveInfo3 -> Alloc(dedebdd6 * 1) nsDiskCacheRecord::RetrieveInfo(7f35f7, 196) mInfoSize=196 mRecordID=104001b mKeyLength=2f nsDiskCacheRecord::RetrieveInfo1 -> Alloc(2f * 1) mMetaDataLength=1000000 nsDiskCacheRecord::RetrieveInfo2 -> Alloc(1000000 * 1) name_len=51020057 nsDiskCacheRecord::RetrieveInfo3 -> Alloc(51020057 * 1) This is on a few gifs on a webpage... It doesn't crash for me (4GB memory in this box) but it allocs all the way up to max datasize ulimit (2G currently). Looks like corrupted data to me, being blindly used in the alloc stuff... Note that I am using the same cache directory for netscape 4.72 as mozilla, could that be triggering stuff? (Which is a case that I think must be handled). My cache.db is 1294336 bytes right now (I have saved a copy of the entire mozilla tree, .mozilla and cache to allow further debugging) so I've been surfing a while with it, but it still shouldn't bail out like this.. This is on Solaris 2.6/Sparc using GCC 2.95.2 with optimize and without debug.
Target Milestone: --- → M17
I am unable to reproduce the crash. Neeti
Target Milestone: M17 → M18
I do see this crash (nsDiskCacheRecord::RetrieveInfo) when my cache fills up (on linux), removing the cache files fixes the crash. This is bad, IMO we definitely don't want this in nsbeta2, this could potentially cause crashes after a while for everyone using nsbeta2. Nominating for nsbeta2.
Keywords: nsbeta2
Putting on [nsbeta2+] radar for beta2 fix.
Whiteboard: [nsbeta2+]
I'm still seeing this now and then. Moving to priority 1 because it is a crasher.
Priority: P3 → P1
I believe this is the crash showing up in talkback data: Windows: http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13387292 http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13361596 http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13508605 http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13478687 http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13403801 Linux: http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13476259 http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13469589 http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13413127 That's 8 of 229 total crashes on a recent list in n.p.m.crash-data. Adding topcrash keyword. If you think these crashes are not related to this bug, please try and figure out what they are! (I think talkback may not always show the top few things in the stack - could it have to do with that it is an optimized build?)
Keywords: topcrash
I´ve once seen a consistent crash in kuro5hin.org... After loading the third page it would crash. Cleaning the cache (via preferences) made it work again. Also, mozilla gets more and more crash-prone the more you fill the cache. Cleaning the cache always make it crash less. At least three of my crashes in kuro5hin.org (which I believe to be cache-related) got sent to talkback (but I have no clue as to how to find them, and I have removed the build while installing a new one, thus losing the talkback ids). They can be recognized by the ´/home/cesarb´ string somewhere in the environment.
Reassigning to neeti
Assignee: gordon → neeti
Status: NEW → ASSIGNED
See also bug 45551 for a related crash where this same code (apparently) blows up the stack.
*** Bug 45772 has been marked as a duplicate of this bug. ***
also seems to be causing crashes on tinderbox and http://www.iplanet.com/downloads/index.html
Whiteboard: [nsbeta2+] → [nsbeta2+] ETA 7/21
Neeti, I can repro this on my Linux box. No luck with NT because of bug 42606
Neeti: Here's quick summary of what I found on this bug. 1> I set MemCache to 3Kb DiskCache to 1Kb 2> close the browser and restart. 3> Before I get the default page to display,I get some odd values in mozilla/netwerk/cache/filecache/nsDiskCacheRecord.cpp in nsDiskCacheRecord::RetrieveInfo(). break at line 418: COPY_INT32(&mRecordID, cur_ptr) ; The value I get for mRecordID looks bogus and consequently triggers the assert at line 435 when mRecordID should be the same as id. My guess is that since I've specified a rediculously small cache size, there is some assumption about the size of the data,mInfoSize, that is picked off just before mRecordID. I hope this helps you track it down. -p
As jst said above, this seems to happen when the cache fills up. I just got into a state where I crashed whenever I tried to load a page (my homepage=about:blank). I crashed 4 times, then I cleared the Memory and Disk caches immediately when I started, and things work fine again.
Hardware: PC → SGI
Changing platform field to "ALL", because it happens on SGI and PC's. We really need checkmark boxes for platforms instead of just one platform or all platforms.
Hardware: SGI → All
reassigning to myself
Assignee: neeti → ruslan
Status: ASSIGNED → NEW
Ok. I got to the bottom of it. The new entry is created all the time, but is only written to the disk when nsHTTPChannel goes away. Therefor if it hits the limit - it will abort and then attempt to evict the entry which is not yet written to the database on disk. Neeti - your proposed kludge is actually a good one; we can even remove the warning from there. The only drawback on itis that it will run DB verification every time this occurs, thus causing degradation in performance. Fixing it will require some considerable changes which at first i started to make, but now I think it's way too much of a change at this point. I'd say go ahead and check it in.
Assignee: ruslan → neeti
I have a fix in hand for this. Neeti
Checked in a fix on the branch and tip
Status: NEW → RESOLVED
Closed: 25 years ago
Resolution: --- → FIXED
What exactly are the performance implications here? Are we doing extra work for every URL that's loaded?
Only when it hits the limit.
We are calling RecoveryCleanup() from nsBrowserInstance in the OnEndDocumentLoad. This does the database recovery every 100 times the cache manager has been accessed. This value can be changed by a preference browser.cache.num.accessed.limit Neeti
Reopening. I'm seeing more page loading problems as a result of cache corruption than I ever was. At one point, my home page refused to load. Trashing the cache fixed it. Why is it that a corrupt cache can halt page loading, anyway? Cache corruption should *never* cause the page to fail to load; it should just cause the page to be refetched from the net.
Status: RESOLVED → REOPENED
Resolution: FIXED → ---
Updating to ETA 7/28
Whiteboard: [nsbeta2+] ETA 7/21 → [nsbeta2+] ETA 7/28
I just pulled with the fix (was it backed out?) and i'm still getting into a corrupted cache state. It doesn't crash for me, i just get to the point after about a day of surfing where no pages load, no error dialog shows up, and the only indication of a problem is a line on the console: "Error loading URL http://bugzilla.mozilla.org/" Ugh.
I'm seeing what pinkerton just described on mac mozilla m18 build 072808 (saw it yesterday on the trunk too) running on OS9. It happens for me after about 6 URLs visited (tested by typing them into the addressbar) URLs tested were from the smoketest page and I seem to fail consistently after loading developer.netscape.com (about the 6th or 7th URL loaded on a fresh install). loading developer.netscape.com as the first URL does not cause the corruption. I can return the browser to a useable state by deleting the contents of the cache folder. This is ogfood for me. It's very nearly a smoketest blocker but there is a workaround, I delete my cache dir contents every 6 or 7 pages visted. I've asked around and people are not reporting this on windows.
Keywords: dogfood
I filed a less serious (but still rather disturbing) crash regression possibly resulting from fixes for this bug as bug 46827. The number of crashes I am seeing on Linux has increased drastically - MTBF is probably around 20-30 mins. I was using the browser for hours without a single crash at the branch point two days ago. However, after I crash once, I start again, normally, and do not need to clear the cache.
Whiteboard: [nsbeta2+] ETA 7/28 → [nsbeta2+][dogfood-] ETA 7/28
marking this dogfood- means i will continue to use 4.x on a daily basis. Having to stop every five or six webpages, quit the app, empty my cache, and relaunch the app (which is painful on mac) is crazy. If this isn't landed by 7/28 (no offense to neeti who may have a patch very close), this better be dogfood+. ....and PDT wonders why no one wants to use 6.0. Geez.
Checked in a fix on branch and tip
Status: REOPENED → RESOLVED
Closed: 25 years ago25 years ago
Resolution: --- → FIXED
Set cache size to 3k and 1k. Ran for about an hour with out any problems. verified: rh6.0 Linux 2000080104
Status: RESOLVED → VERIFIED
*** Bug 46827 has been marked as a duplicate of this bug. ***
tip build from 8/25/00. it's baaaaaaaack. had to toss my cache so i could load any more webpages. reopening. again, no crash, but cache is corrupt to the point where i can no longer load _and_ page.
Status: VERIFIED → REOPENED
Keywords: nsbeta3
Resolution: FIXED → ---
*** This bug has been marked as a duplicate of 50977 ***
Status: REOPENED → RESOLVED
Closed: 25 years ago25 years ago
Resolution: --- → DUPLICATE
Status: RESOLVED → VERIFIED
verif. DUP
You need to log in before you can comment on or make changes to this bug.