Closed Bug 476649 Opened 17 years ago Closed 16 years ago

Fix/Handle |Connection to the other side was lost in a non-clean fashion.| on TB tinderboxes

Categories

(Mozilla Messaging Graveyard :: Server Operations, defect)

defect
Not set
normal

Tracking

(Not tracked)

RESOLVED FIXED

People

(Reporter: sgautherie, Assigned: gozer)

References

Details

Per bug 471083 comment 15. Issue still occurring: http://tinderbox.mozilla.org/showlog.cgi?log=Thunderbird/1233670487.1233670947.12168.gz Linux comm-central trunk build on 2009/02/03 06:14:47
Component: Build Config → Server Operations
Product: Thunderbird → Mozilla Messaging
QA Contact: build-config → server-ops
Version: Trunk → other
For some reason, that builder didn't get his keepalive disabled, but reduced to 60 seconds. I tried that originally in trying to figure out what would make a difference. Must have forgotten it behind when converting all buildbots to keepalive = None. Disabled keepalives now there too, should not happen again.
Status: NEW → RESOLVED
Closed: 17 years ago
Resolution: --- → FIXED
http://tinderbox.mozilla.org/showlog.cgi?log=Thunderbird3.0/1234460993.1234461500.22975.gz Win2k3 comm-central bloat build on 2009/02/12 09:49:53 (still) had this failure...
(In reply to comment #2) And http://tinderbox.mozilla.org/showlog.cgi?log=Thunderbird/1234460370.1234461523.23020.gz Win2k3 comm-central mozilla-central bloat build on 2009/02/12 09:39:30 http://tinderbox.mozilla.org/showlog.cgi?log=Thunderbird/1234460758.1234461504.22985.gz Win2k3 comm-central trunk build on 2009/02/12 09:45:58 too.
This really is not (fully) fixed: http://tinderbox.mozilla.org/showlog.cgi?log=Thunderbird/1234760191.1234761223.17375.gz Win2k3 comm-central trunk build on 2009/02/15 20:56:31 http://tinderbox.mozilla.org/showlog.cgi?log=Thunderbird/1234760191.1234761213.17362.gz Win2k3 comm-central mozilla-central bloat build on 2009/02/15 20:56:31 http://tinderbox.mozilla.org/showlog.cgi?log=Thunderbird/1234782000.1234785057.7748.gz Win2k3 comm-central trunk nightly on 2009/02/16 03:00:00 http://tinderbox.mozilla.org/showlog.cgi?log=Thunderbird/1234794912.1234795454.16518.gz Win2k3 comm-central trunk build on 2009/02/16 06:35:12 http://tinderbox.mozilla.org/showlog.cgi?log=Thunderbird/1234794937.1234795393.16404.gz Win2k3 comm-central check on 2009/02/16 06:35:37 http://tinderbox.mozilla.org/showlog.cgi?log=Thunderbird3.0/1234794754.1234795454.16521.gz Win2k3 comm-central bloat build on 2009/02/16 06:32:34
Status: RESOLVED → REOPENED
Resolution: FIXED → ---
(In reply to comment #5) And http://tinderbox.mozilla.org/showlog.cgi?log=Thunderbird/1234793739.1234795510.16675.gz Win2k3 comm-central mozilla-central bloat build on 2009/02/16 06:15:39
I've been looking at logs for the potential source of this, and there is one thing that has somewhat jumped to the front. buildbot (on the server side of things) generate *lots* of logs and rotates them when they reach over a certain fixed size (100k by default), but it does so in what appears to be a somewhat interestingly brain-damaged way. you've got : [...] twisted.log.3 twisted.log.2 teisted.log.1 twisted.log And when it's time to rotate logs, it does mv twisted.log.3 twisted.log.4 mv twisted.log.2 twisted.log.3 mv twisted.log.1 twisted.log.2 mv twisted.log twisted.log.1 etc, and it seems to do so in a blocking loop where nothing else can be hapenning. I mention this because most of these disconnecting errors show up at the beginining of one of these log files. So, what I am suspecting happens, is that when there are *thousands* of such log files, the log /rotation/ step will end up taking quite a bit of time, and possibly trip up liveness detection of slaves in some way. I've found an alternate twisted logging implementation that rotates log files daily, and timestamps them, so I am going to try that on the unittest buildbot master. application = service.Application('buildmaster') +from twisted.python.log import ILogObserver, FileLogObserver +from twisted.python.logfile import DailyLogFile +logfile = DailyLogFile("twistd.log", "logs") +application.setComponent(ILogObserver, FileLogObserver(logfile).emit) BuildMaster('.', configfile).setServiceParent(application)
(In reply to comment #8) > http://tinderbox.mozilla.org/showlog.cgi?log=Thunderbird3.0/1234800382.1234802524.6898.gz > Linux comm-1.9.1 check on 2009/02/16 08:06:22 This one and the Linux comm-central check were caused by me when I restarted the slaves, so unrelated.
I've moved some VMs around to create more CPU/IO space for the buildbot master as of Feb 24 07:34 PDT, so I'd like to watch and see if more of these disconnects happen after that time.
I've not seen any of these for a while, except in the cases that we know that buildbot has been restarted or network connectivity lost. I think we've pretty much solved the core issue here - gozer: want to resolve?
Assignee: nobody → gozer
Status: REOPENED → RESOLVED
Closed: 17 years ago → 16 years ago
Resolution: --- → FIXED
You need to log in before you can comment on or make changes to this bug.