Closed
Bug 40084
Opened 26 years ago
Closed 25 years ago
[CRASH] Crash in disk cache code
Categories
(Core :: Networking: Cache, defect, P1)
Tracking
()
M18
People
(Reporter: jce2, Assigned: neeti)
References
()
Details
(Keywords: crash, topcrash, Whiteboard: [nsbeta2+][dogfood-] ETA 7/28)
Core was generated by `./mozilla-bin'.
Program terminated with signal 6, Aborted.
(gdb) backtrace
backtrace
#0 0x405e551d in __mmap () from /lib/libc.so.6
#1 0x40595e4d in chunk_alloc (ar_ptr=0x40620a00, nb=1701406320) at malloc.c:1809
#2 0x40595972 in __libc_malloc (bytes=1701406313) at malloc.c:2596
#3 0x4026aacc in PR_Malloc (size=1701406313) at prmem.c:38
#4 0x4011baac in nsAllocatorImpl::Alloc (this=0x80663e8, size=1701406313) at
nsAllocator.cpp:82
#5 0x4011bb7a in nsAllocator::Alloc (size=1701406313) at nsAllocator.cpp:148
#6 0x40dee049 in nsDiskCacheRecord::RetrieveInfo (this=0x95a8528,
aInfo=0x8742fb0, aInfoLength=386) at nsDiskCacheRecord.cpp:442
#7 0x40deacf5 in nsDBEnumerator::GetNext (this=0x95a7ba8, _retval=0xbffff260)
at nsDBEnumerator.cpp:100
#8 0x40de7a8a in nsReplacementPolicy::AddAllRecordsInCache (this=0x866cb98,
aCache=0x8609128) at nsReplacementPolicy.cpp:170
#9 0x40de8974 in nsReplacementPolicy::LoadAllRecordsInAllCacheDatabases
(this=0x866cb98) at nsReplacementPolicy.cpp:619
#10 0x40de89e9 in nsReplacementPolicy::Evict (this=0x866cb98,
aTargetOccupancy=4849) at nsReplacementPolicy.cpp:646
#11 0x40de2969 in nsCacheManager::LimitDiskCacheSize () at nsCacheManager.cpp:516
#12 0x40de299b in nsCacheManager::LimitCacheSize () at nsCacheManager.cpp:526
#13 0x40de98dd in CacheOutputStream::Write (this=0x8b026c0, aBuf=0x9375800
"-article.cgi?article_key=1000351120\">2. oh......</A>(aquiri, a.k.a. extreme
burger, no happy meals here / 04.19.00)<DD>\n<DL>\n<DT><A
HREF=\"read-article.cgi?article_key=1000351176\">1. You aren't doing y"...,
aCount=1460, aActualBytes=0xbffff374) at nsCacheEntryChannel.cpp:89
#14 0x40de6b34 in InterceptStreamListener::write (this=0x918fa38, aBuf=0x9375800
"-article.cgi?article_key=1000351120\">2. oh......</A>(aquiri, a.k.a. extreme
burger, no happy meals here / 04.19.00)<DD>\n<DL>\n<DT><A
HREF=\"read-article.cgi?article_key=1000351176\">1. You aren't doing y"...,
aNumBytes=1460) at nsCachedNetData.cpp:1192
#15 0x40de6bc8 in InterceptStreamListener::Read (this=0x918fa38, buf=0x9375800
"-article.cgi?article_key=1000351120\">2. oh......</A>(aquiri, a.k.a. extreme
burger, no happy meals here / 04.19.00)<DD>\n<DL>\n<DT><A
HREF=\"read-article.cgi?article_key=1000351176\">1. You aren't doing y"...,
count=1460, aActualBytes=0xbffff490) at nsCachedNetData.cpp:1180
#16 0x417bd796 in nsParser::OnDataAvailable (this=0x93382e0, channel=0x8e153b8,
aContext=0x0, pIStream=0x918fa3c, sourceOffset=0, aLength=1460) at nsParser.cpp:1840
#17 0x40ee7a70 in nsDocumentOpenInfo::OnDataAvailable (this=0x8e15130,
aChannel=0x8e153b8, aCtxt=0x0, inStr=0x918fa3c, sourceOffset=0, count=1460) at
nsURILoader.cpp:189
#18 0x40e16350 in nsHTTPFinalListener::OnDataAvailable (this=0x8e151a8,
aChannel=0x8e153b8, aContext=0x0, aStream=0x918fa3c, aSourceOffset=0,
aCount=1460) at nsHTTPResponseListener.cpp:1216
#19 0x40de6d2f in InterceptStreamListener::OnDataAvailable (this=0x918fa38,
channel=0x8e153b8, ctxt=0x0, inStr=0x91ddd5c, sourceOffset=0, count=1460) at
nsCachedNetData.cpp:1164
#20 0x40e1486d in nsHTTPServerListener::OnDataAvailable (this=0x9280700,
channel=0x8de5964, context=0x8e153b8, i_pStream=0x91ddd5c,
i_SourceOffset=471892, i_Length=1460) at nsHTTPResponseListener.cpp:540
#21 0x40db0e34 in nsOnDataAvailableEvent::HandleEvent (this=0x95a83d8) at
nsAsyncStreamListener.cpp:406
#22 0x40db0056 in nsStreamListenerEvent::HandlePLEvent (aEvent=0x95a8400) at
nsAsyncStreamListener.cpp:97
#23 0x40110e93 in PL_HandleEvent (self=0x95a8400) at plevent.c:575
#24 0x40110d2d in PL_ProcessPendingEvents (self=0x811bd28) at plevent.c:520
#25 0x40112a91 in nsEventQueueImpl::ProcessPendingEvents (this=0x8129330) at
nsEventQueue.cpp:316
#26 0x40a6344b in event_processor_callback (data=0x8129330, source=6,
condition=GDK_INPUT_READ) at nsAppShell.cpp:143
#27 0x40a6306a in our_gdk_io_invoke (source=0x81b47d8, condition=G_IO_IN,
data=0x81cd040) at nsAppShell.cpp:56
#28 0x40c1d0c7 in g_io_unix_dispatch (source_data=0x81b4838,
current_time=0xbffff884, user_data=0x81cd040) at giounix.c:135
#29 0x40c1e6c7 in g_main_dispatch (current_time=0xbffff884) at gmain.c:652
#30 0x40c1ec7b in g_main_iterate (block=1, dispatch=1) at gmain.c:870
#31 0x40c1ee09 in g_main_run (loop=0x81c0308) at gmain.c:928
#32 0x40b4aa47 in gtk_main () at gtkmain.c:475
#33 0x40a63a9a in nsAppShell::Run (this=0x811b9d8) at nsAppShell.cpp:313
#34 0x40700edf in nsAppShellService::Run (this=0x8122608) at
nsAppShellService.cpp:371
#35 0x805319b in main1 (argc=1, argv=0xbffffb54, nativeApp=0x0) at
nsAppRunner.cpp:904
#36 0x80538c0 in main (argc=1, argv=0xbffffb54) at nsAppRunner.cpp:1188
#37 0x40556ab2 in __libc_start_main (main=0x80536f4 <main>, argc=1,
argv=0xbffffb54, init=0x804eda8 <_init>, fini=0x805cf88 <_fini>,
rtld_fini=0x4000b0b0 <_dl_fini>, stack_end=0xbffffb4c) at
../sysdeps/generic/libc-start.c:78
Comment 1•26 years ago
|
||
jce2@po.cwru.edu, could you please explain how to reproduce the bug?
Comment 3•26 years ago
|
||
Recommend resolving as WORKSFORME, due to lack of detail.
| Reporter | ||
Comment 4•26 years ago
|
||
This is a SPORATIC bug..it's somewhat difficult to reproduce.
Try loading long pages and clicking on one of the links while it is still
downloading. Many times, the browser will just freeze up and then crash.
jce2@po.cwru.edu, can you give a url of a long page where I can reproduce the
crash?
Neeti
Comment 6•26 years ago
|
||
I have been able to track down some of this
Here is some debug output:
First one that looks ok:
nsDiskCacheRecord::RetrieveInfo(7f20b8, 1a6)
mInfoSize=1a6
mRecordID=10a0082
mKeyLength=31
nsDiskCacheRecord::RetrieveInfo1 -> Alloc(31 * 1)
mMetaDataLength=13b
nsDiskCacheRecord::RetrieveInfo2 -> Alloc(13b * 1)
name_len=26
nsDiskCacheRecord::RetrieveInfo3 -> Alloc(26 * 1)
nsDiskCacheRecord::GetMetaData -> Alloc(13b * 1)
Here are a few not-so-good ones..
nsDiskCacheRecord::RetrieveInfo(7f2c68, 195)
mInfoSize=195
mRecordID=1070004
mKeyLength=44
nsDiskCacheRecord::RetrieveInfo1 -> Alloc(44 * 1)
mMetaDataLength=31303a31
nsDiskCacheRecord::RetrieveInfo2 -> Alloc(31303a31 * 1)
name_len=0
nsDiskCacheRecord::RetrieveInfo3 -> Alloc(0 * 1)
nsDiskCacheRecord::RetrieveInfo(7f1d69, 132)
mInfoSize=132
mRecordID=1060046
mKeyLength=26
nsDiskCacheRecord::RetrieveInfo1 -> Alloc(26 * 1)
mMetaDataLength=21094039
nsDiskCacheRecord::RetrieveInfo2 -> Alloc(21094039 * 1)
name_len=dedebdd6
nsDiskCacheRecord::RetrieveInfo3 -> Alloc(dedebdd6 * 1)
nsDiskCacheRecord::RetrieveInfo(7f35f7, 196)
mInfoSize=196
mRecordID=104001b
mKeyLength=2f
nsDiskCacheRecord::RetrieveInfo1 -> Alloc(2f * 1)
mMetaDataLength=1000000
nsDiskCacheRecord::RetrieveInfo2 -> Alloc(1000000 * 1)
name_len=51020057
nsDiskCacheRecord::RetrieveInfo3 -> Alloc(51020057 * 1)
This is on a few gifs on a webpage...
It doesn't crash for me (4GB memory in this box) but it allocs all the way up
to max datasize ulimit (2G currently).
Looks like corrupted data to me, being blindly used in the alloc stuff...
Note that I am using the same cache directory for netscape 4.72 as mozilla,
could that be triggering stuff? (Which is a case that I think must be handled).
My cache.db is 1294336 bytes right now (I have saved a copy of the entire
mozilla tree, .mozilla and cache to allow further debugging) so I've been
surfing a while with it, but it still shouldn't bail out like this..
This is on Solaris 2.6/Sparc using GCC 2.95.2 with optimize and without debug.
I am unable to reproduce the crash.
Neeti
Target Milestone: M17 → M18
Comment 8•26 years ago
|
||
I do see this crash (nsDiskCacheRecord::RetrieveInfo) when my cache fills up (on
linux), removing the cache files fixes the crash.
This is bad, IMO we definitely don't want this in nsbeta2, this could
potentially cause crashes after a while for everyone using nsbeta2.
Nominating for nsbeta2.
Keywords: nsbeta2
| Reporter | ||
Comment 10•26 years ago
|
||
I'm still seeing this now and then. Moving to priority 1 because it is a
crasher.
Priority: P3 → P1
I believe this is the crash showing up in talkback data:
Windows:
http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13387292
http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13361596
http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13508605
http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13478687
http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13403801
Linux:
http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13476259
http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13469589
http://cyclone/reports/stackcommentemail.cfm?dynamicBBID=13413127
That's 8 of 229 total crashes on a recent list in n.p.m.crash-data. Adding
topcrash keyword. If you think these crashes are not related to this bug,
please try and figure out what they are! (I think talkback may not always show
the top few things in the stack - could it have to do with that it is an
optimized build?)
Keywords: topcrash
Comment 12•26 years ago
|
||
I´ve once seen a consistent crash in kuro5hin.org... After loading the third
page it would crash. Cleaning the cache (via preferences) made it work again.
Also, mozilla gets more and more crash-prone the more you fill the cache.
Cleaning the cache always make it crash less.
At least three of my crashes in kuro5hin.org (which I believe to be
cache-related) got sent to talkback (but I have no clue as to how to find them,
and I have removed the build while installing a new one, thus losing the
talkback ids). They can be recognized by the ´/home/cesarb´ string somewhere in
the environment.
Comment 14•25 years ago
|
||
See also bug 45551 for a related crash where this same code (apparently) blows
up the stack.
| Assignee | ||
Comment 15•25 years ago
|
||
*** Bug 45772 has been marked as a duplicate of this bug. ***
Comment 16•25 years ago
|
||
also seems to be causing crashes on tinderbox and
http://www.iplanet.com/downloads/index.html
Comment 17•25 years ago
|
||
Neeti, I can repro this on my Linux box. No luck with NT because of bug 42606
Comment 18•25 years ago
|
||
Neeti:
Here's quick summary of what I found on this bug.
1> I set MemCache to 3Kb
DiskCache to 1Kb
2> close the browser and restart.
3> Before I get the default page to display,I get some odd values in
mozilla/netwerk/cache/filecache/nsDiskCacheRecord.cpp
in nsDiskCacheRecord::RetrieveInfo().
break at line 418: COPY_INT32(&mRecordID, cur_ptr) ;
The value I get for mRecordID looks bogus and consequently
triggers the assert at line 435 when mRecordID should be the
same as id.
My guess is that since I've specified a rediculously small cache
size, there is some assumption about the size of the data,mInfoSize, that is
picked off just before mRecordID.
I hope this helps you track it down.
-p
As jst said above, this seems to happen when the cache fills up. I just got
into a state where I crashed whenever I tried to load a page (my
homepage=about:blank). I crashed 4 times, then I cleared the Memory and Disk
caches immediately when I started, and things work fine again.
| Reporter | ||
Comment 20•25 years ago
|
||
Changing platform field to "ALL", because it happens on SGI and PC's.
We really need checkmark boxes for platforms instead of just one platform or all
platforms.
Hardware: SGI → All
Comment 22•25 years ago
|
||
Ok. I got to the bottom of it. The new entry is created all the time, but is
only written to the disk when nsHTTPChannel goes away. Therefor if it hits the
limit - it will abort and then attempt to evict the entry which is not yet
written to the database on disk. Neeti - your proposed kludge is actually a good
one; we can even remove the warning from there. The only drawback on itis that
it will run DB verification every time this occurs, thus causing degradation in
performance. Fixing it will require some considerable changes which at first i
started to make, but now I think it's way too much of a change at this point.
I'd say go ahead and check it in.
Assignee: ruslan → neeti
| Assignee | ||
Comment 23•25 years ago
|
||
I have a fix in hand for this.
Neeti
| Assignee | ||
Comment 24•25 years ago
|
||
Checked in a fix on the branch and tip
Status: NEW → RESOLVED
Closed: 25 years ago
Resolution: --- → FIXED
Comment 25•25 years ago
|
||
What exactly are the performance implications here? Are we doing extra work for
every URL that's loaded?
Comment 26•25 years ago
|
||
Only when it hits the limit.
| Assignee | ||
Comment 27•25 years ago
|
||
We are calling RecoveryCleanup() from nsBrowserInstance in the
OnEndDocumentLoad. This does the database recovery every 100 times the cache
manager has been accessed. This value can be changed by a preference
browser.cache.num.accessed.limit
Neeti
Comment 28•25 years ago
|
||
Reopening. I'm seeing more page loading problems as a result of cache corruption
than I ever was. At one point, my home page refused to load. Trashing the cache
fixed it.
Why is it that a corrupt cache can halt page loading, anyway? Cache corruption
should *never* cause the page to fail to load; it should just cause the page to
be refetched from the net.
Status: RESOLVED → REOPENED
Resolution: FIXED → ---
Comment 30•25 years ago
|
||
I just pulled with the fix (was it backed out?) and i'm still getting into a
corrupted cache state. It doesn't crash for me, i just get to the point after
about a day of surfing where no pages load, no error dialog shows up, and the
only indication of a problem is a line on the console:
"Error loading URL http://bugzilla.mozilla.org/"
Ugh.
Comment 31•25 years ago
|
||
I'm seeing what pinkerton just described on mac mozilla m18 build 072808 (saw it
yesterday on the trunk too) running on OS9. It happens for me after about 6
URLs visited (tested by typing them into the addressbar) URLs tested were from
the smoketest page and I seem to fail consistently after loading
developer.netscape.com (about the 6th or 7th URL loaded on a fresh install).
loading developer.netscape.com as the first URL does not cause the corruption.
I can return the browser to a useable state by deleting the contents of the
cache folder. This is ogfood for me. It's very nearly a smoketest blocker but
there is a workaround, I delete my cache dir contents every 6 or 7 pages visted.
I've asked around and people are not reporting this on windows.
Keywords: dogfood
I filed a less serious (but still rather disturbing) crash regression possibly
resulting from fixes for this bug as bug 46827. The number of crashes I am
seeing on Linux has increased drastically - MTBF is probably around 20-30 mins.
I was using the browser for hours without a single crash at the branch point two
days ago. However, after I crash once, I start again, normally, and do not need
to clear the cache.
Comment 33•25 years ago
|
||
marking this dogfood- means i will continue to use 4.x on a daily basis. Having
to stop every five or six webpages, quit the app, empty my cache, and relaunch
the app (which is painful on mac) is crazy.
If this isn't landed by 7/28 (no offense to neeti who may have a patch very
close), this better be dogfood+.
....and PDT wonders why no one wants to use 6.0. Geez.
| Assignee | ||
Comment 34•25 years ago
|
||
Checked in a fix on branch and tip
Status: REOPENED → RESOLVED
Closed: 25 years ago → 25 years ago
Resolution: --- → FIXED
Comment 35•25 years ago
|
||
Set cache size to 3k and 1k. Ran for about an hour with out any problems.
verified:
rh6.0 Linux 2000080104
Status: RESOLVED → VERIFIED
| Assignee | ||
Comment 36•25 years ago
|
||
*** Bug 46827 has been marked as a duplicate of this bug. ***
Comment 37•25 years ago
|
||
tip build from 8/25/00.
it's baaaaaaaack. had to toss my cache so i could load any more webpages.
reopening. again, no crash, but cache is corrupt to the point where i can no
longer load _and_ page.
pinkerton: Was what you describe bug 49108?
| Assignee | ||
Comment 39•25 years ago
|
||
*** This bug has been marked as a duplicate of 50977 ***
Status: REOPENED → RESOLVED
Closed: 25 years ago → 25 years ago
Resolution: --- → DUPLICATE
Updated•25 years ago
|
Status: RESOLVED → VERIFIED
Comment 40•25 years ago
|
||
verif. DUP
You need to log in
before you can comment on or make changes to this bug.
Description
•