Closed Bug 700672 (r4-dongles) Opened 14 years ago Closed 14 years ago

figure out what's going on with talos-r4 dongles

Categories

(Infrastructure & Operations :: RelOps: General, task)

x86
macOS
task
Not set
major

Tracking

(Not tracked)

RESOLVED FIXED

People

(Reporter: jhford, Assigned: jhford)

References

()

Details

Attachments

(3 files)

We are having lots of issues with dongles on talos-r4-snow machines (haven't been looking for them in talos-r4-lion machines yet). There are at least two failure cases: 1) machines boot, but resolution is set to 800x600x32 after puppet runs but before test harnesses run 2) machine doesn't think it has a dongle, fails to puppet. (i verified that these machines don't have 1600x1200 available) talos-r4-snow-015.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-029.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-030.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-031.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-032.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-034.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-038.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-039.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-040.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-042.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-046.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-050.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-053.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-054.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-058.build.mozilla.org Display 0: 1280x1024x32 I think we need to figure out why this is happening and fix it post-haste. Some things I think we should try: A) putting 600Ω resistor across hsync and vsync [1] B) putting 75Ω resistors across the RGB pins instead of 100Ω [1] C) use a DVI dongle [2] [1] http://www.mp3car.com/vbulletin/general-hardware-discussion/1763-how-does-a-video-card-detect-vga-monitor-is-present.html#post11453 [2] http://www.monoprice.com/products/product.asp?c_id=101&cp_id=10110&cs_id=1011003&p_id=3048&seq=1&format=2
Blocks: 693918
Anyone know if something happened to the slaves listed in comment 1? Was any work done near these machines?
Assignee: nobody → server-ops-releng
Component: Release Engineering → Server Operations: RelEng
QA Contact: release → zandr
Asked jhford where he was tracking failure frequency, etc, and he didn't have a source other than this bug at present. We *should* start tracking that data though. jhford: can we extract some loose approximation of the failures/frequency from the rev4 bugs and comments that philor has filed?
No of us have been in that datacenter this week, so I don't think anyone bumped into them. Do you have a list of which hosts fall into which category? Having the right resolution before puppet seems to indicate a software issue. The machine thinking it doesn't have a dongle at all sounds more like it may be a hardware issue, though. It sounds like there may be multiple issues? Getting a list of the error conditions of each host (and date stamped each time it occurs) would be very helpful as would knowing if these hosts have ever worked, worked once, etc. The more data we have, the easier it will hopefully be to find patterns.
I have gone through all the bugs that I can find. Of interest is that twice, machines that were intermittently failing eventually started persistently failing The following machines have issues: 011,015,018,029,030,031,032,034,038,039,040,046,052,053,054,058,061,069,074 - have failed in the persistent failure state at least once (some were fixed by reseating dongles) 023,033,070,078,080 - have failed in the intermittent state at least once 042,050 - have failed both in the intermittent state *and* persistent state 057,064 - have failed at least twice in the intermittent state I think we need to look closer at the dongles. Apparently, we are using 100 Ohm resistors across the RGB signals. I wonder if the new motherboard's sensor circuit thinks 100 Ohm is right on the threshold of having no monitor present. Lets try using a different dongle type as well as a DVI monitor simulator. The new dongle type has the following pin out: 1 & 6 connected by 75 Ohm resistor 2 & 7 connected by 75 Ohm resistor 3 & 8 connected by 75 Ohm resistor 10 & 13 connected by 600 Ohm resistor 10 & 14 connected by 600 Ohm resistor The DVI monitor simulator is http://www.monoprice.com/products/product.asp?c_id=101&cp_id=10110&cs_id=1011003&p_id=3048&seq=1&format=2 Please install the new dongle on 042,057,023,011,015 Please install the dvi simulator on 050,064,033,018,029 using an HDMI->DVI converter and program the dvi simulators using a Dell Ultrasharp U2410 raw data: https://docs.google.com/spreadsheet/ccc?key=0AlguvtTDx79WdFZVVVk3RFJ3TzNDS0d1aHN1bE5Tdmc
Severity: normal → major
Assignee: server-ops-releng → arich
I'll work on this - I'll get the monoprice item delivered to scl1. I'll order the DVI simulator from monoprice, and a dozen more solder cups - http://search.digikey.com/scripts/DkSearch/dksus.dll?Detail&name=17EHD-015-P-AA-0-00-ND Jake, how are your soldering skills?
do we already have the 75Ω and 600Ω resistors?
I know I am going to regret admitting this but my solder skills are excellent.
All of these are now showing up with 1600x1200. I think it would be a good idea to check the screen resolution on all of the r4 machines at this point, because there was a lot of crawling around and trying to reseat dongles and push in cables. It's very *very* tight in these racks, so I want to make sure we haven't dislodged anything else. Some of the ones below were touching metal chassis or other metal bits of the rack, so that could easily account for a short that would temporarily take out the dongle. This isn't the case for all of them, though. Since we keep getting asked to fish around in back of these machines, my concern is that we're fixing some and then breaking others in the process (all it takes is bumping two of them together, or bumping one into a rack, I bet) since they're really fragile.
We now have new machines with dongle problems: talos-r4-snow-062.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-065.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-067.build.mozilla.org Display 0: 800x600x32 talos-r4-snow-070.build.mozilla.org Display 0: 1280x1024x32
(In reply to John Ford [:jhford] from comment #10) > We now have new machines with dongle problems: to be clear, these are the only machines not currently at the right resolution
Almost forgot to check the lion machines. I think this issue is affecting them too. I'd like to concentrate on the snow leopard machines. As this is looking like a hardware issue, I think that we should fix we do for the snow machines should be done for lion as well. The following lion machines are currently having dongle issues: talos-r4-lion-049.build.mozilla.org Display 0: 1280x1024x32 talos-r4-lion-051.build.mozilla.org Display 0: 1280x1024x32 talos-r4-lion-062.build.mozilla.org Display 0: 1280x1024x32
Summary: figure out what's going on with talos-r4-snow dongles → figure out what's going on with talos-r4 dongles
Blocks: 696417
Blocks: 696453
Blocks: 695679
I can't find 600W resistors that have a reasonable tolerance. Here's what I'm looking at: 1 12 17EHD-015-P-AA-0-00-ND CONN HD DB15 MALE GOLD FLASH 0 1.89000 $22.68 2 50 75ADCT-ND RES 75 OHM 1/4W 0.1% MF AXL 0 0.26460 $13.23 3 50 604XBK-ND RES 604 OHM 1/4W 1% METAL FILM 0 0.09100 $4.55 I'm assuming that the 75 ohm resistors are probably the tighter tolerance, and the 600's are pull-up/pull-down's, so 604 += 6.04 (1%) is OK. I haven't ordered these yet, though, so please speak up if this is incorrect.
I'll reseat these tomorrow: talos-r4-lion-049.build.mozilla.org Display 0: 1280x1024x32 talos-r4-lion-051.build.mozilla.org Display 0: 1280x1024x32 talos-r4-lion-062.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-062.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-065.build.mozilla.org Display 0: 1280x1024x32 talos-r4-snow-067.build.mozilla.org Display 0: 800x600x32 talos-r4-snow-070.build.mozilla.org Display 0: 1280x1024x32
colo-trip: --- → scl1
one example calling for 75 ohm resistors: http://bitclockers.com/forums/index.php?topic=7.0 Not sure where the 600 ohm's figure in.
colo-trip: scl1 → ---
The DVI Doctor (monoprice) and soldering stuff (comment 13) are ordered, for delivery to 650 castro attn matt larrain.
(In reply to Dustin J. Mitchell [:dustin] from comment #13) > I can't find 600W resistors that have a reasonable tolerance. 1% is a reasonable tolerance. 0.1% is crazy exotic precision stuff, which is why they're crazy expensive. (26 cents for a resistor is crazy expensive) > I'm assuming that the 75 ohm resistors are probably the tighter tolerance, > and the 600's are pull-up/pull-down's, so 604 += 6.04 (1%) is OK. I haven't > ordered these yet, though, so please speak up if this is incorrect. All these are doing is simulating the impedance of the monitor inputs. Nothing we're doing is emulating DDC/EDID, so at best the macs will think there's "something" out there, but won't know what it is. The RGB video signals expect a 75ohm impedance. I can't find a solid reference for the input impedance for H-sync and V-sync, but "I read it on the internet" class references say ~2k. 600 ohms is rather a lot of load for a driver expecting 2k. Further, I'd be very surprised if anyone designed a system to look for load on such a high-impedance input. High impedance means there's very little load on the driver, which makes detecting said load very difficult. Add in impedance mismatches due to the various connectors (VGA connectors are ~100ohm impedance, can't even begin to speculate about MiniDP) and this becomes a pretty fuzzy way to detect a monitor. tl;dr going down to 50-75ohms on the RGB lines may help. I don't expect the resistors on Hsync/Vsync to matter. DVIdoctors will work, but are expensive and annoying to configure. At least they don't need power except for programming.
Great info, thanks. Yeah, $0.26 each is a *lot*, but I figured if 100 ohms is out of range, then the tolerances are fairly tight, so 5% would still leave room for ambiguity. Jake mentioned he's used 68Ω before. From this info, it sounds like the things to try are: * replace 100-ohm with 75-ohm dongles on five minis, bake for a few days, and * add the DVI Doctor to one node if the DVI Doctor doesn't work, we should look for software problems. if the DVI Doctor works, but the 75-ohm do not, * add 600Ω resistors to h-sync/v-sync on those cups, and * try the same with 68-ohm resistors, with and without h-sync/v-sync. I'd be happy to start with 2kΩ instead, although sources I see around the 'net are using 470Ω and 1kΩ, so maybe that's flexible. Let's hope it doesn't get there. Jake: solder cups with different resistances will be easily mixed up, so if you can find some easy way to label them, or at least make them quickly visually distinguishable, that'd be great.
digi-key order is http://www.fedex.com/Tracking?language=english&cntry_code=us&tracknumbers=0126%200607%201554%20715 (not shipped yet, but was shipped ground so ~end of the week)
These were grouped pretty physically close to each other: talos-r4-lion-049 1280x1024x16 1280x1024x32 1600x1200x8 1600x1200x16 talos-r4-lion-051 1280x1024x16 1280x1024x32 1600x1200x8 1600x1200x16 talos-r4-lion-062 1280x1024x16 1280x1024x32 1600x1200x8 1600x1200x16 talos-r4-snow-062 1600x1200x8 1600x1200x16 1600x1200x32 1600x1200x8 talos-r4-snow-065 1600x1200x8 1600x1200x16 1600x1200x32 1600x1200x8 talos-r4-snow-067 1600x1200x8 1600x1200x16 1600x1200x32 1600x1200x8 talos-r4-snow-070 1600x1200x8 1600x1200x16 1600x1200x32 1600x1200x8 We wound up putting in some repaired hardware, so it would be interesting to know if others failed near where we were working.
(In reply to Amy Rich [:arich] [:arr] from comment #20) > These were grouped pretty physically close to each other: They all currently show as being at 1600x1200x32, are these machines that had failures in the spreadsheet?
Those were the ones that you said needed to be reseated in this bug, so I did so. I just did a grep for 1600x1200|head -1.
(In reply to John Ford [:jhford] from comment #5) > I have gone through all the bugs that I can find. Of interest is that > twice, machines that were intermittently failing eventually started > persistently failing > raw data: > https://docs.google.com/spreadsheet/ > ccc?key=0AlguvtTDx79WdFZVVVk3RFJ3TzNDS0d1aHN1bE5Tdmc Is there a key that says what means what for this data? I didn't see any more entries after yesterday, so does that mean that no hosts are currently showing errors? I'm presuming that puppet is notifying you when the resolution is incorrect, right, so you should know fairly quickly when the permanent failure state happens and a reboot won't fix it.
(In reply to Amy Rich [:arich] [:arr] from comment #22) > Those were the ones that you said needed to be reseated in this bug, so I > did so. I just did a grep for 1600x1200|head -1. ahhh, that makes sense :) (In reply to Amy Rich [:arich] [:arr] from comment #23) > (In reply to John Ford [:jhford] from comment #5) > > I have gone through all the bugs that I can find. Of interest is that > > twice, machines that were intermittently failing eventually started > > persistently failing > > > raw data: > > https://docs.google.com/spreadsheet/ > > ccc?key=0AlguvtTDx79WdFZVVVk3RFJ3TzNDS0d1aHN1bE5Tdmc > > Is there a key that says what means what for this data? its not a well organized spreadsheet, but its a list of machines. each time there is a failure, the bug number is put in the next available column to the right. For machines that don't get stuck in a failing state, a date/time stamp is added as well as the bug number > I didn't see any more entries after yesterday, so does that mean that no > hosts are currently showing errors? it means that I haven't updated the spreadsheet > I'm presuming that puppet is notifying you when the resolution is incorrect, > right, so you should know fairly quickly when the permanent failure state > happens and a reboot won't fix it. puppet notifying us? we don't have any of that, sadly.
Going to hand this off to jake since I'm leaving town tonight and he'll be doing the soldering work.
Assignee: arich → jwatkins
I am recommending we seal all the dongles with an Acrylic Conformal Coating. I've used this stuff before to protect pcb prototypes and it works great. It has a fast cure time that doesn't require heat and its viscosity can be adjusted. http://www.mgchemicals.com/products/419b.html http://www.mgchemicals.com/downloads/pdf/specsheets/419l.pdf
An alternative option to the acrylic: install housing on each connector. http://www.allelectronics.com/make-a-store/item/DB-15H/DB-15-HOOD/1.html
Summarizing, we have three things going on here: * Jake is working on ways of insulating the dongles * Resistors and whatnot (comment 13) have arrived in mtv1 - on Matt's desk * DVI Doctor has arrived on 11/16 - on Matt's desk Jake's got the helm as far as the insulation goes. For the experimental new dongles, the next steps are outlined in comment #18, and subject to revision as Jake sees fit.
colo-trip: --- → scl1
(In reply to Dustin J. Mitchell [:dustin] from comment #28) > Summarizing, we have three things going on here: > > * Jake is working on ways of insulating the dongles > * Resistors and whatnot (comment 13) have arrived in mtv1 - on Matt's desk > * DVI Doctor has arrived on 11/16 - on Matt's desk > > Jake's got the helm as far as the insulation goes. For the experimental new > dongles, the next steps are outlined in comment #18, and subject to revision > as Jake sees fit. How are things progressing? Please let me know when these tests are ready so the machines can be taken out of production for installation and testing.
John, we'll be ready for these on arriving at scl1 tomorrow morning. Can you disable and post a list of slaves that we should put these on? Jake has 5 dongles, plus the DVI Doctor, so we'll need six hosts.
Assignee: jwatkins → jhford
I guess we aren't going to do 5 simulators as requested in comment 5 Please put the simulator on 042 and a new dongle on 050, 064,033,018,029. I've disabled these slaves in slavealloc.
(In reply to John Ford [:jhford] from comment #31) > I guess we aren't going to do 5 simulators as requested in comment 5 The discussion progressed significantly from there, given a lot of new information, and we're working off the more recent plan in comment 18. I didn't get any feedback to the contrary. Are the numbers in comment 31 snow's or lions?
comment 31 is referring to snow leopard machines
(In reply to John Ford [:jhford] from comment #31) > I guess we aren't going to do 5 simulators as requested in comment 5 > > Please put the simulator on 042 and a new dongle on 050, 064,033,018,029. > I've disabled these slaves in slavealloc. I have installed the dvi Dr on snow-042 and the new 75ohm dongles on snow-050, 064,033,018 and 029.
All machines in comment 34 have been re-enabled in slavealloc.
(In reply to John Ford [:jhford] from comment #35) > All machines in comment 34 have been re-enabled in slavealloc. Any update on these machines and associated failures, John?
i am letting these run in production for a while to build up data.
Well, it looks like we are seeing something odd happening in bug 709436 to talos-r4-snow-064. The failure mode is changing, where the machines are reporting 1600x1200x32 but are failing with symptoms of being at 800x600x32. This is not encouraging for the new dongle design as implemented. Lets go back to the original plan of having multiple dvi-doctors. Can we overnight 9 additional dvi-doctors to SCL1 to be installed on 057,023,011,015,050,064,033,018,029?
From what I can tell, this sounds like a software error - either in 'screenresolution', or in the tests themselves. If those are giving different results, then one of them is incorrect, and we should stop trusting that result. Also, it doesn't look like https://docs.google.com/spreadsheet/ccc?key=0AlguvtTDx79WdFZVVVk3RFJ3TzNDS0d1aHN1bE5Tdmc#gid=0 has been kept up to date, which makes it hard to see what patterns there might be here. DVI Doctors are useful for experimentation and data-gathering, but will not work well in production - they are large and heavy and will most likely fall off of the connectors or pull them out of the minis. We can get more DVI Doctors on Monday, but given the above (likely software problem, incomplete data, and not a production solution), can you give some details on how that will help to figure out what's wrong here?
DVI Doctors are ordered: Order Number: 5491824 - Your Order Detail The courier price ("Overnight Express") was cheapest, so I went with that option. They don't ship on weekends, so that means it will be over-Monday-night, and arrive in mtv1 on Tuesday. Jake, let's plan to get these programmed in mtv1 on Tuesday afternoon, after they arrive.
(In reply to Dustin J. Mitchell [:dustin] from comment #39) > From what I can tell, this sounds like a software error - either in > 'screenresolution', or in the tests themselves. If those are giving > different results, then one of them is incorrect, and we should stop > trusting that result. Is that the only possible explanation? Since my Mac changes resolution multiple times per day as I plug in and unplug the external monitor, I would have said the most likely explanation for screenresolution saying one thing at the start of every run, and then sometimes it being the same and sometimes it being different after n minutes when we hit mochitest-chrome, and sometimes it being the same as it was for screenresolution after m minutes in mochitest-browser-chrome and other times it being the same as it was for mochitest-chrome and yet other times being different than it was for either screenresolution or for mochitest-chrome, would be that the resolution is changing during the course of the run.
It's by no means the only explanation. I'm hoping we can get together and share data on what's gone wrong, and what possible solutions we have.
If we're going to gather data about how the dongles are working, we probably need some way of keeping those slaves working. If buildapi/recent/ is telling the truth about last jobs, https://build.mozilla.org/buildapi/recent/talos-r4-snow-018 - 2011-12-07 22:44 https://build.mozilla.org/buildapi/recent/talos-r4-snow-029 - current https://build.mozilla.org/buildapi/recent/talos-r4-snow-033 - 2011-12-07 23:04 https://build.mozilla.org/buildapi/recent/talos-r4-snow-042 - current (after a two day layoff ending last night) https://build.mozilla.org/buildapi/recent/talos-r4-snow-050 - 2011-12-04 09:49 https://build.mozilla.org/buildapi/recent/talos-r4-snow-064 - current (though it's blowing so many jobs it probably shouldn't be, bug 709436)
Good point. If the talos-r4-snow pool is so under-utilized that the slaves sit idle for days, maybe we should disable half (or some other fraction) of them temporarily?
It was late, but I think jhford said that 042 had lost its connection to its master after a reconfig, and he rebooted it - if I heard him right, 018 and 033 are probably in the same state (and maybe even 050 from a previous incident).
(In reply to Dustin J. Mitchell [:dustin] from comment #39) > From what I can tell, this sounds like a software error - either in > 'screenresolution', or in the tests themselves. If those are giving > different results, then one of them is incorrect, and we should stop > trusting that result. > > Also, it doesn't look like > https://docs.google.com/spreadsheet/ > ccc?key=0AlguvtTDx79WdFZVVVk3RFJ3TzNDS0d1aHN1bE5Tdmc#gid=0 has been kept up > to date, which makes it hard to see what patterns there might be here. I put that together not as the authoritative data source, but as something for me to work with to get numbers. > DVI Doctors are useful for experimentation and data-gathering, but will not > work well in production - they are large and heavy and will most likely fall > off of the connectors or pull them out of the minis. I think it is too early to say that they will not work in production. DVI Doctors can use the HDMI->DVI dongle that is already included in R4 and R5 minis. We have shown that the minidp connectors do not reliably make a good connection. The HDMI cables are significantly more solid. There is shelf space behind each mini, I think we can find a way to mount this box to each shelf. If DVI Doctors present a clear solution to this problem, I think we ought to consider them the long term fix. I don't want to spend yet more time finding the cheapest solution if we have a solid solution. > We can get more DVI Doctors on Monday, but given the above (likely software > problem, incomplete data, and not a production solution), can you give some > details on how that will help to figure out what's wrong here? software vs. hardware is not an important distinction. What is import to differentiate is whether we have control to fix the defect. If the bug is in the display apis, driver, OS or hardware, there is nothing we can do to fix it and must work around it. I don't think we can call the dvi doctor "not a production solution" yet. It sounds a lot more like a 'production' solution than a minidp plug that consistently does not make a solid connection. (In reply to Phil Ringnalda (:philor) from comment #45) > It was late, but I think jhford said that 042 had lost its connection to its > master after a reconfig, and he rebooted it - if I heard him right, 018 and > 033 are probably in the same state (and maybe even 050 from a previous > incident). Yes, 042 had lost its connection to the master, probably a network 'blip'. It has been back in the pool and has not experienced any dongle related issues so far, though, its too early to tell if the dvi doctor fixes the problem. (In reply to Dustin J. Mitchell [:dustin] from comment #44) > Good point. If the talos-r4-snow pool is so under-utilized that the slaves > sit idle for days, maybe we should disable half (or some other fraction) of > them temporarily? More likely, if they aren't reaching production masters, they are stuck in the failure state of having the wrong resolution.
(In reply to John Ford [:jhford] from comment #46) > If DVI Doctors present a clear solution to this problem, I think we ought to > consider them the long term fix. I don't want to spend yet more time > finding the cheapest solution if we have a solid solution. That may be fair - let's see how the next 10 work out. > More likely, if they aren't reaching production masters, they are stuck in > the failure state of having the wrong resolution. Correct -- bug 709596. At this point, that bug is blocking gathering additional data on this problem, because any host that comes up with an incorrect resolution hangs there until manually fixed.
I have a new version of screenresolution which prints more information. The logs are visible on stderr as well as the standard apple logging mechanism. Each run has at least the argv printed to logs, so we can see exactly how many times the program is being invoked and when. ~/software/screenresolution $ ./screenresolution 2011-12-13 11:27:44.945 screenresolution[19469:707] starting screenresolution argv=./screenresolution 2011-12-13 11:27:44.946 screenresolution[19469:707] Incorrect command line ~/software/screenresolution $ ./screenresolution get 2011-12-13 11:27:49.953 screenresolution[19471:707] starting screenresolution argv=./screenresolution get 2011-12-13 11:27:49.958 screenresolution[19471:707] Display 0: 1920x1200x32 ~/software/screenresolution $ ./screenresolution set 2011-12-13 11:27:52.305 screenresolution[19472:707] starting screenresolution argv=./screenresolution set ~/software/screenresolution $ ./screenresolution set 1600x1200x32 2011-12-13 11:27:57.833 screenresolution[19473:707] starting screenresolution argv=./screenresolution set 1600x1200x32 2011-12-13 11:27:57.837 screenresolution[19473:707] set mode on display 0 to 1600x1200x32 ~/software/screenresolution $ ./screenresolution list 2011-12-13 11:28:06.904 screenresolution[19476:707] starting screenresolution argv=./screenresolution list Available Modes on Display 0 1920x1200x16 1920x1200x32 1920x1200x30 960x600x16 960x600x32 960x600x30 1680x1050x16 1680x1050x32 1680x1050x30 1600x1200x16 1600x1200x32 1600x1200x30 1600x1200x16 1600x1200x32 1600x1200x30 1280x1024x16 1280x1024x32 1280x1024x30 1280x1024x16 1280x1024x32 1280x1024x30 1152x720x16 1152x720x32 1152x720x30 1024x768x16 1024x768x32 1024x768x30 1024x768x16 1024x768x32 1024x768x30 1024x640x16 1024x640x32 1024x640x30 1280x800x16 1280x800x32 1280x800x30 800x600x16 800x600x32 800x600x30 800x600x16 800x600x32 800x600x30 800x500x16 800x500x32 800x500x30 640x480x16 640x480x32 640x480x30 640x480x16 640x480x32 640x480x30 720x480x16 720x480x32 720x480x30 720x480x16 720x480x32 720x480x30 1280x960x16 1280x960x32 1280x960x30 1280x960x16 1280x960x32 1280x960x30 1344x1008x16 1344x1008x32 1344x1008x30 1344x840x16 1344x840x32 1344x840x30 1600x1000x16 1600x1000x32 1600x1000x30
Attachment #581351 - Flags: review?(coop)
Jake has the dongles and will get them programmed today for installation on Thursday.
(In reply to Dustin J. Mitchell [:dustin] from comment #49) > Jake has the dongles and will get them programmed today for installation on > Thursday. This doesn't leave us with a week before the meeting on Dec 21. Is it possible to install them today or tomorrow instead?
Nobody's onsite until Thursday. We'll still have 6 days', which should give us something to guess on -- even if that means we agree to share the data next Thursday and discuss in IRC.
Alias: r4-dongles
As a point of order, the DVI doctor we currently have installed is connected via a standard DVI cable and a HDMI->DVI adapter. It could also use a MiniDP->DVI adapter, although my own experience with the video on these devices suggests that the two outputs are not symmetrical, so that may change the results. All of the dongles with solder cups are MiniDP->VGA, so if we choose to go with DVI Doctors, we will need more adapters as well, and probably some shorter DVI cables.
HDMI->DVI cables were included with every rev4 mini we bought [1]. Lets reuse the ones we already own. My preference is that we use the HDMI->DVI converter because we already have them, because they are official apple parts and because they seem to have a significantly more solid connection than mini-dp cables. We could also try out these: http://www.monoprice.com/products/product.asp?c_id=102&cp_id=10231&cs_id=1023104&p_id=2661&seq=1&format=2 if we decide to go with DVI Doctors. [1] http://support.apple.com/kb/SP585
Yeah, Zandr pointed that out right after I posted the comment.
(In reply to John Ford [:jhford] from comment #46) > I think it is too early to say that they will not work in production. It is also too early to say that they will. I'd like to get to a DC and look at how they're installed, but I agree with Dustin that they are mechanically challenging. They are also labor intensive to program, and we don't know anything about their failure modes yet. > DVI Doctors can use the HDMI->DVI dongle that is already included in R4 and R5 > minis. We have shown that the minidp connectors do not reliably make a good > connection. The HDMI cables are significantly more solid. "Shown" implies that you have data. If you have data that shows that the HDMI connection is more reliable than the Mini-DP, then you haven't shared it. In order to control for the connector, we'd have to be using the same solution (DVI Doctor or resistor dongle) on both connectors. As HDMI has no analog, that means we'd need to test with DVI Doctors on Mini-DP to run that experiment. > There is shelf space behind each mini, I think we can find a way to mount > this box to each shelf. I'm glad you think so. I think those shelves are already fairly full of cables, and probably don't make a good mounting location. Someone who has actually installed a DVI Doctor in a full rack should comment here. I'm particularly concerned about servicablility (which is already pretty poor) once we get a number of them packed together. Remember also that since both the HDMI adapter and the DVI-Doctor have female connectors, we need a short DVI cable between them. This is starting to be a lot of hardware to cram into that shelf, and we'll probably have to engineer some sort of mounting rail inside the rack. Were the machines that are getting the DVI Doctors tomorrow selected for adjacency? Could we do that, and figure out if we can actually install these rationally on every unit? > If DVI Doctors present a clear solution to this problem, I think we ought to > consider them the long term fix. I don't want to spend yet more time > finding the cheapest solution if we have a solid solution. I think we need to evaluate data from all of the solutions before we start jumping to conclusions, and weigh that against the operational considerations of each solution. There are costs to each solution that go beyond releng's time and the bill from Monoprice. > software vs. hardware is not an important distinction. Uhh, what? If you are declining to identify the problem, I find it very difficult to understand how you expect to come up with a solution. > I don't think we can call the dvi doctor "not a production solution" yet. > It sounds a lot more like a 'production' solution than a minidp plug that > consistently does not make a solid connection. I don't think we can call *anything* a production solution yet. I do think it's very easy for you to discount the operational concerns because they aren't your problem. As I said, there are costs to each solution that go beyond your time and the cost of the DVI doctors. > More likely, if they aren't reaching production masters, they are stuck in > the failure state of having the wrong resolution. Do we have logging that could determine that (either before or after the improved tool in comment 48)? There's too much speculation and not enough data here, let's fix that.
Attachment #581351 - Flags: review?(coop) → review+
I will need a list of slaves to attach the 9 new DVI Doctors. John, can you provide this please? I will be on site at SCL1 today.
057,023,011,015,050,064,033,018,029 (that's from comment 38, but there's been a lot of churn since then)
(In reply to John Ford [:jhford] from comment #57) > 057,023,011,015,050,064,033,018,029 As I asked in comment #55: > Were the machines that are getting the DVI Doctors tomorrow selected for > adjacency? Could we do that, and figure out if we can actually install these > rationally on every unit? Yes, no, maybe?
9 DVI Doctors have been installed on 057,023,011,015,050,064,033,018,029.
(In reply to Zandr Milewski [:zandr] from comment #55) > (In reply to John Ford [:jhford] from comment #46) > > I think it is too early to say that they will not work in production. > > It is also too early to say that they will. I'd like to get to a DC and look > at how they're installed, but I agree with Dustin that they are mechanically > challenging. They are also labor intensive to program, and we don't know > anything about their failure modes yet. Yes, that's why we are trying to gather data here. I wouldn't want to go through the trouble of installing them on 160 machines to find out that they don't work. > > DVI Doctors can use the HDMI->DVI dongle that is already included in R4 and R5 > > minis. We have shown that the minidp connectors do not reliably make a good > > connection. The HDMI cables are significantly more solid. > > "Shown" implies that you have data. If you have data that shows that the > HDMI connection is more reliable than the Mini-DP, then you haven't shared > it. In order to control for the connector, we'd have to be using the same > solution (DVI Doctor or resistor dongle) on both connectors. As HDMI has no > analog, that means we'd need to test with DVI Doctors on Mini-DP to run that > experiment. Matt and I debugged this with multiple computers. It turns out that even though the dongles were fully inserted, they weren't being picked up by the computer. Matt tried pushing the dongle in to make sure it was fully inserted. The dongle still wasn't being detected by the computer. The only fix was for Matt to completely remove and reinsert the dongle. This affected a bunch of the slaves, multiple times. To clarify, both my aggregation and the raw data are public. > > There is shelf space behind each mini, I think we can find a way to mount > > this box to each shelf. > > I'm glad you think so. I think those shelves are already fairly full of > cables, and probably don't make a good mounting location. Someone who has > actually installed a DVI Doctor in a full rack should comment here. I'm > particularly concerned about servicablility (which is already pretty poor) > once we get a number of them packed together. > > Remember also that since both the HDMI adapter and the DVI-Doctor have > female connectors, we need a short DVI cable between them. This is starting > to be a lot of hardware to cram into that shelf, and we'll probably have to > engineer some sort of mounting rail inside the rack. As I mentioned in comment 53, we could look at HDMI->DVI cables. > Were the machines that are getting the DVI Doctors tomorrow selected for > adjacency? Could we do that, and figure out if we can actually install these > rationally on every unit? They were selected for having demonstrated the issues we are trying to correct for. I still think this is the correct basis for selection. > > If DVI Doctors present a clear solution to this problem, I think we ought to > > consider them the long term fix. I don't want to spend yet more time > > finding the cheapest solution if we have a solid solution. > > I think we need to evaluate data from all of the solutions before we start > jumping to conclusions, and weigh that against the operational > considerations of each solution. There are costs to each solution that go > beyond releng's time and the bill from Monoprice. Yes, I have been asking to get the DVI Doctors installed since November 9 so we could start collecting data on possible fixes. I am also aware that there are other costs. > > software vs. hardware is not an important distinction. > > Uhh, what? If you are declining to identify the problem, I find it very > difficult to understand how you expect to come up with a solution. I believe what I said was: (comment 46) > software vs. hardware is not an important distinction. What is import to > differentiate is whether we have control to fix the defect. If the bug is > in the display apis, driver, OS or hardware, there is nothing we can do to > fix it and must work around it. I don't see anything there that suggests that I am uninterested in figuring out what is causing these problems. In fact, I think finding out what is happening is so important that I filed this very bug to figure out whats going on with the dongles! What I am trying to say is that if the problem turns out to be something outside of our control, we need to work around it. > > I don't think we can call the dvi doctor "not a production solution" yet. > > It sounds a lot more like a 'production' solution than a minidp plug that > > consistently does not make a solid connection. > > I don't think we can call *anything* a production solution yet. I do think > it's very easy for you to discount the operational concerns because they > aren't your problem. As I said, there are costs to each solution that go > beyond your time and the cost of the DVI doctors. I think its pretty obvious that any solution is going to have costs in excess of acquisition. I think we also need to keep in mind the opportunity cost of working on this problem instead of others as well as the cost of wasted developer time dealing with fallout from this issue. I don't appreciate your characterization and don't think its correct. I have jumped in many times to help out with things that aren't "my problem". I also don't think that this is relevant to the discussion at hand. > > More likely, if they aren't reaching production masters, they are stuck in > > the failure state of having the wrong resolution. > > Do we have logging that could determine that (either before or after the > improved tool in comment 48)? The improvements in comment 48 concern logging in the actual application. We don't have logging to check when the machines get into the failing state. > There's too much speculation and not enough data here, let's fix that. Getting data to see what's going wrong is the purpose of this bug. What you are referring to as speculation, I can only assume is people coming up with ideas and trying to implement a test for their idea. If you have a solution, please, do share it!
I'm curious is all r4 machines are exhibiting these problems, or just a select number (and if those issues are ongoing. John, are you tracking each occurrence now?). If there are machines that are not exhibiting the problems, are we attaching the same hardware solutions we're testing to a control group of functioning machines?
(In reply to John Ford [:jhford] from comment #60) > Matt and I debugged this with multiple computers. It turns out that even > though the dongles were fully inserted, they weren't being picked up by the > computer. Matt tried pushing the dongle in to make sure it was fully > inserted. The dongle still wasn't being detected by the computer. The only > fix was for Matt to completely remove and reinsert the dongle. This > affected a bunch of the slaves, multiple times. This might be the most important thing that has been said in this bug, and I think we've all been misinterpreting that information. To me, this doesn't say anything about connector reliability, but rather indicates that a state change (unplugging and replugging the dongle) is required to resolve the failure. Beyond that, we have no data on HDMI connector reliability, so we can't make any assertions about the relative reliability of HDMI vs. MiniDP. > To clarify, both my aggregation and the raw data are public. In comment 46, you indicated that the aggregation was not authoritative. Is there an authoritative aggregation? > As I mentioned in comment 53, we could look at HDMI->DVI cables. Yup, and I haven't found a source for those shorter than 3'. You aren't going to get two 3' cables and two DVI doctors onto a shelf that's already full of several cables. > They were selected for having demonstrated the issues we are trying to > correct for. I still think this is the correct basis for selection. Fair enough. > I don't appreciate your characterization and don't think its correct. I > have jumped in many times to help out with things that aren't "my problem". > I also don't think that this is relevant to the discussion at hand. It's quite relevant. It is very easy to ignore the costs of any solution that are outside your organization. Every group in this organization is guilty of that to some degree. But taking a step back and thinking about it, the operational ramifications of DVI Doctors far outweigh some additional debugging effort on the resistor-based solution. We think that we're having trouble with connector reliability, so we're going to add more connectors in a quest to fix it? From an operations standpoint that's the wrong direction. > > > More likely, if they aren't reaching production masters, they are stuck in > > > the failure state of having the wrong resolution. > > > > Do we have logging that could determine that (either before or after the > > improved tool in comment 48)? > > The improvements in comment 48 concern logging in the actual application. > We don't have logging to check when the machines get into the failing state. Do we have any evidence that supports the assertion that machines that aren't talking to the production masters are in that state because of a resolution problem? That's the gap in my understanding. > Getting data to see what's going wrong is the purpose of this bug. What you > are referring to as speculation, I can only assume is people coming up with > ideas and trying to implement a test for their idea. If you have a > solution, please, do share it! I've been pointing out experimental design issues. I just want to be very careful that we are supporting the conclusions we draw, which I've not seen in this bug so far.
(In reply to Zandr Milewski [:zandr] from comment #62) > (In reply to John Ford [:jhford] from comment #60) > > > The only > > fix was for Matt to completely remove and reinsert the dongle. This > > affected a bunch of the slaves, multiple times. > > This might be the most important thing that has been said in this bug, and I > think we've all been misinterpreting that information. > > To me, this doesn't say anything about connector reliability, but rather > indicates that a state change (unplugging and replugging the dongle) is > required to resolve the failure. Thinking a bit more about this. I believe that this means the connector works *fine*. This raises a couple of questions, and suggests an experiment. Question: Does this error condition survive a reboot? If so, we should try the following experiment: If it isn't already, set "restart after power failure" (systemsetup -setrestartpowerfailure on). IMO, this should be set anyway. Wait for a machine to enter the 'bad' state, then: Shut down the machine (not a restart) Do a PDU reset. Let the machine come back up. If that fixes the problem, then we know that a warm-reset is not resetting the parts of the gfx controller that are hosed, and a cold-reset does. I'm not a fan of doing PDU resets since it slows down the reboot cycle, but I'd take it over DVI doctors.
(In reply to Zandr Milewski [:zandr] from comment #63) > Shut down the machine (not a restart) > > Do a PDU reset. Let the machine come back up. A little further reading suggests this might not work, that the restart-after-power-cut only works if the system was on when it lost power. Testing this while onsite would be a good thing. :D
Depends on: 711374
Depends on: 711376
Depends on: 711382
(In reply to Amy Rich [:arich] [:arr] from comment #61) > I'm curious is all r4 machines are exhibiting these problems, or just a > select number (and if those issues are ongoing. John, are you tracking each > occurrence now?). If there are machines that are not exhibiting the > problems, are we attaching the same hardware solutions we're testing to a > control group of functioning machines? There isn't much historic data on machines that don't make it to buildbot right now. Intermittent failures of machines that have made it past the puppet check for resolution are being tracked in bugs 693918, 695679, 696417 and 696453. I have filed: bug 711374 for talos-r4-snow-018 bug 711376 for talos-r4-snow-012 bug 711382 for reseating dongles on 42 machines. In the process for gathering this data, I have found a new failure mode. The quartz window server is failing, but ssh continues to work. In this mode, I can VNC into the machine but the VNC screen is solid blue with a working cursor and nothing else. One machine initially had broken sshd along with the blue screen, but eventually allowed me to log in. This failure mode seems to recover after a reboot initiated over ssh. I saw this mode on machines with working dongles, with malfunctioning dongles and with dvi doctors. I've created an etherpad to track instances of the consistent failure state with data starting today. Updating it is a mostly manual task (automatically generated data that requires human processing and verification) and takes a while to do. https://etherpad.mozilla.org/rev4-dongle-log
bug 711382 looks like it contains a bunch of machines that haven't been seen before. I only have the spreadsheet linked in comment 5, but from that data none of the lion machines and only 6 of 22 snow machines are repeat offenders. This suggests that the answer to :arr's question in comment 61 is 'no'. This appears to be across the entire pool, and not restricted to certain machines. None of the machines in bug 711382 have 75ohm dongles, but without historical data on the failures, it's hard to evaluate that. I also don't understand the dependency on bug 711374 and bug 711376 here. The procedure requested in those bugs will not produce any data about resolution or dongle state.
(In reply to John Ford [:jhford] from comment #65) > https://etherpad.mozilla.org/rev4-dongle-log Is there anything to indicate that modes D and E are in any way related to dongles or screen resolution? I could see that mode F is the result of resolutions changing and the window server getting confused.
(In reply to Zandr Milewski [:zandr] from comment #67) > (In reply to John Ford [:jhford] from comment #65) > > https://etherpad.mozilla.org/rev4-dongle-log > > Is there anything to indicate that modes D and E are in any way related to > dongles or screen resolution? not specifically, just logging as much info as I can and these are fairly common cases > I could see that mode F is the result of resolutions changing and the window > server getting confused. Yah, I think that's very reasonable, especially considering that the screens go blue like that during mode changes. (In reply to Zandr Milewski [:zandr] from comment #66) > I also don't understand the dependency on bug 711374 and bug 711376 here. > The procedure requested in those bugs will not produce any data about > resolution or dongle state. There may or may not be something wrong with the dongles on these machines. I am happy to unlink the bugs if it turns out that the dongle is not the cause of problems on this machine.
I just found out that we *replaced* the 75ohm dongles with DVI-Doctors, ending the 75ohm experiment. I don't believe that we have sufficient data to draw any conclusions, therefore you should not take the next sentence seriously, but... No mini with a 75ohm resistor dongle has failed. I would reocmmend selecting some repeat offenders and swapping in 75ohm dongles.
(In reply to John Ford [:jhford] from comment #68) > There may or may not be something wrong with the dongles on these machines. > I am happy to unlink the bugs if it turns out that the dongle is not the > cause of problems on this machine. Nothing in those bugs says anything about dongles one way or the other. I would recommend unlinking them.
(In reply to Zandr Milewski [:zandr] from comment #69) > No mini with a 75ohm resistor dongle has failed. talos-r4-snow-064?
(In reply to Phil Ringnalda (:philor) from comment #71) > talos-r4-snow-064? Oh, good catch. snow-064 shows as Mode B in the etherpad. Hm. I do wish we had 75ohm dongles still in service.
It occurs to me I have some relevant anecdotal evidence from home. I have an r5 mini with two screens, and on wake from sleep in some significant proportion of cases, it will only detect one of the screens. "Detect Displays" will generally fix this. I wonder if there's a way to trigger "Detect Displays" from software? That might provide a path to a solution here.
Haven't found anything that's part of the distribution, but I did find this: http://code.google.com/p/detectdisplays/
(In reply to Zandr Milewski [:zandr] from comment #74) > Haven't found anything that's part of the distribution, but I did find this: > http://code.google.com/p/detectdisplays/ tried that and it didn't seem to change anything: talos-r4-lion-012:~ cltbld$ screenresolution list Available Modes on Display 0 1x1x8 1x1x16 1x1x32 1x1x64 1x1x96 800x600x8 800x600x16 800x600x32 800x600x64 800x600x96 1024x768x8 1024x768x16 1024x768x32 1024x768x64 1024x768x96 1280x1024x8 1280x1024x16 1280x1024x32 1280x1024x64 1280x1024x96 1680x1050x8 1680x1050x16 1680x1050x32 1680x1050x64 1680x1050x96 1280x1024x32 talos-r4-lion-012:~ cltbld$ ./detectdisplays talos-r4-lion-012:~ cltbld$ screenresolution list Available Modes on Display 0 1x1x8 1x1x16 1x1x32 1x1x64 1x1x96 800x600x8 800x600x16 800x600x32 800x600x64 800x600x96 1024x768x8 1024x768x16 1024x768x32 1024x768x64 1024x768x96 1280x1024x8 1280x1024x16 1280x1024x32 1280x1024x64 1280x1024x96 1680x1050x8 1680x1050x16 1680x1050x32 1680x1050x64 1680x1050x96 1280x1024x32
(In reply to Zandr Milewski [:zandr] from comment #72) > (In reply to Phil Ringnalda (:philor) from comment #71) > > > talos-r4-snow-064? > > Oh, good catch. snow-064 shows as Mode B in the etherpad. > > Hm. I do wish we had 75ohm dongles still in service. See also: bug 709436. This slave exhibited all sorts of new problems on the 75Ω dongle.
https://bugzilla.mozilla.org/show_bug.cgi?id=711382#c9 is relevant. Cold starts recover machines from this state without touching the dongle.
Changing from an onlyif to unless means that we are explicitly search for a string instead of searching for its absence. This should be generally better. This also redirects the output from NSLog from stderr to stdout for grep to read. The relevant sections of the puppet type manifest are: -onlyif: If this parameter is set, then this exec will only run if the command returns 0. -unless: If this parameter is set, then this exec will run unless the command returns 0
Attachment #582977 - Flags: review?
Attachment #582977 - Flags: review? → review?(aki)
Attachment #582977 - Flags: review?(aki) → review+
Depends on: 712180
(In reply to Zandr Milewski [:zandr] from comment #77) > https://bugzilla.mozilla.org/show_bug.cgi?id=711382#c9 is relevant. > > Cold starts recover machines from this state without touching the dongle. Moving discussion over here, as that bug will go away when the machines are operational again. (In reply to Zandr Milewski [:zandr] from comment #9) > (In reply to John Ford [:jhford] from comment #8) > > > Yes, it recovered but is now broken again. Its last test was at "Fri Dec 16 > > 14:36:44 2011" (pst). This test was successful and at the correct > > resolution. > > Excellent. We now know two more things: > > 1) It has nothing to do with connectors on the dongles. Had the fix stuck, I'd completely agree with you. Given that the machine broke again so soon, I wouldn't personally rule this out as part of the problem. It seems we are right on the edge of what causes the machine not to detect a display. > 2) Unpleasant though it may be, a cold reboot will recover a machine. > > I'll play around with pmset to see if I can come up with a solution that > will let us do a shutdown instead of a reboot after each test. Yep, looks like pmset has 'acwake' and 'autorestart' which look like they are both useful. There is also 'shutdown -u' which halts the machine but leaves the power on so the power loss setting properly turns the machine on when the pdu gives it power. That said, I think we should consider ethernet PDUs + pmset + shutdown -u a last resort if the 75Ω dongles and DVI Doctor don't work. A solution like this adds a lot of complexity to automate it and a lot of ongoing human-time if we keep it manual.
(In reply to John Ford [:jhford] from comment #79) > Had the fix stuck, I'd completely agree with you. Given that the machine > broke again so soon, I wouldn't personally rule this out as part of the > problem. I don't see that this has anything to do with the connectors. We recovered 8 out of 8 machines without touching the connector, and we've established previously that wiggling the connectors doesn't help, only removing and reinstalling them helps. That's the same sort of state change we're going through by doing a power cycle. > It seems we are right on the edge of what causes the machine not > to detect a display. This seems to be the case, but the machines we power cycled all had 100Ω dongles, right? > Yep, looks like pmset has 'acwake' and 'autorestart' which look like they > are both useful. There is also 'shutdown -u' which halts the machine but > leaves the power on so the power loss setting properly turns the machine on > when the pdu gives it power. I wasn't going that direction, actually. 'acwake' doesn't apply, and as you say, autorestart would require shutdown -u and a power cycle. More in a bit. > That said, I think we should consider ethernet PDUs + pmset + shutdown -u a > last resort if the 75Ω dongles and DVI Doctor don't work. I had been under the impression you'd ruled out 75Ω dongles, glad that's not the case. I really want to rule out DVI Doctors. ;) > A solution like > this adds a lot of complexity to automate it and a lot of ongoing human-time > if we keep it manual. Using PDU's sure. But there's another way: pmset schedule wakeorpoweron "`date -v+1M "+%m/%d/%y %H:%M:%S"`" && shutdown -h now This sets an 'alarm' to power on the machine at now +1 minute and then does a shutdown. I believe that this is equivalent to the "Test 1" in bug 711382, doing a clean shutdown and then pressing the power button.
(In reply to Zandr Milewski [:zandr] from comment #80) > (In reply to John Ford [:jhford] from comment #79) > > It seems we are right on the edge of what causes the machine not > > to detect a display. > > This seems to be the case, but the machines we power cycled all had 100Ω > dongles, right? Yep, they were the 100Ω ones. > Using PDU's sure. But there's another way: > > pmset schedule wakeorpoweron "`date -v+1M "+%m/%d/%y %H:%M:%S"`" && shutdown > -h now > > This sets an 'alarm' to power on the machine at now +1 minute and then does > a shutdown. > > I believe that this is equivalent to the "Test 1" in bug 711382, doing a > clean shutdown and then pressing the power button. Can you test this? That looks neat. What happens if the machine takes longer than 1 minute to shut down? Can we set more than one alarm time? My reading of the manpage suggests we can. If that's the case, we could schedule a wakeorpoweron for 1, 5 and 60 minutes. Before we move to this, we should have a machine continually reboot by this method to see how often or if it fails.
talos-r4-snow-013:~ cltbld$ screenresolution list 2011-12-20 07:53:42.407 screenresolution[1325:903] starting screenresolution argv=screenresolution list Available Modes on Display 0 1x1x8 1x1x16 1x1x32 1x1x64 1x1x96 800x600x8 800x600x16 800x600x32 800x600x64 800x600x96 1024x768x8 1024x768x16 1024x768x32 1024x768x64 1024x768x96 1280x1024x8 1280x1024x16 1280x1024x32 1280x1024x64 1280x1024x96 1680x1050x8 1680x1050x16 1680x1050x32 1680x1050x64 1680x1050x96 1280x1024x32 talos-r4-snow-013:~ cltbld$ screenresolution get 2011-12-20 07:53:51.752 screenresolution[1326:903] starting screenresolution argv=screenresolution get 2011-12-20 07:53:51.759 screenresolution[1326:903] Display 0: 1280x1024x32 <hedule wakeorpoweron "`date -v+1M "+%m/%d/%y %H:%M:%S"`" && shutdown -h nowConnection to talos-r4-snow-013.build.mozilla.org closed by remote host. Connection to talos-r4-snow-013.build.mozilla.org closed. Last login: Tue Dec 20 07:54:34 2011 from bm-vpn01.build.sjc1.mozilla.com talos-r4-snow-013:~ cltbld$ screenresolution get 2011-12-20 07:54:56.168 screenresolution[123:903] starting screenresolution argv=screenresolution get 2011-12-20 07:54:56.194 screenresolution[123:903] Display 0: 1280x1024x32 talos-r4-snow-013:~ cltbld$ screenresolution list 2011-12-20 07:54:59.824 screenresolution[146:903] starting screenresolution argv=screenresolution list Available Modes on Display 0 1x1x8 1x1x16 1x1x32 1x1x64 1x1x96 800x600x8 800x600x16 800x600x32 800x600x64 800x600x96 1024x768x8 1024x768x16 1024x768x32 1024x768x64 1024x768x96 1280x1024x8 1280x1024x16 1280x1024x32 1280x1024x64 1280x1024x96 1680x1050x8 1680x1050x16 1680x1050x32 1680x1050x64 1680x1050x96 1280x1024x32 talos-r4-snow-013:~ cltbld$ Sadly, it looks like this doesn't solve the problem.
(In reply to John Ford [:jhford] from comment #82) > Sadly, it looks like this doesn't solve the problem. but.... using a 3m time worked! I think its a capacitor or inductor holding charge between power cycles.
The corresponding script has the contents: ~/mozilla/new-r4-reboot $ cat pmset-reboot.sh #!/bin/bash # Schedule a boot at 3, 10 and 60 minutes in case machine takes too long to # shutdown. # Clear out any jobs left over from previous boots. Doing this here as a # safety net. We also don't know if there are adverse effects to having # a wakeorpoweron event firing on a system in production rm /Library/Preferences/SystemConfiguration/com.apple.AutoWake.plist for i in 3 10 60 ; do pmset schedule wakeorpoweron "$(date -v+${i}M '+%m/%d/%y %H:%M:%S')" done This will also need a change to /etc/sudoers to allow passwordless root on pmset-reboot.sh and to disallow passwordless '/sbin/reboot'. We'll also need to update tools to call pmset-reboot.sh instead of reboot.
Attachment #583255 - Flags: review?(coop)
(In reply to John Ford [:jhford] from comment #84) > Created attachment 583255 [details] [diff] [review] > changes to puppet to deploy the new reboot script This is what's needed to test out the pmset scheduled boot.
Here's the etherpad gathering everything I can find on the topic so far (excluding raw data, but linking heavily): https://etherpad.mozilla.org/2011-12-21-dongles-mtng
Depends on: 712750
John -- do you want to schedule another meeting sometime next week to see where we're at with this? I can do the scheduling if you're prefer, but I thought I'd share the fun :)
I have not seen any intermittent or consistent failure from any of the 75Ω dongle or dvi doctor machines. talos-r4-snow-015 has died, but it is unpingable and nothing suggests that this is a dongle or resolution issue. We were seeing intermittent problems roughly once per day and have a multiple machines fail in the consistent state each day. I think we can call the 75Ω dongles and dvi doctors both a solution and we can say that the problem was that we had the incorrect resistance. I would expect that in the two weeks we've had the 75Ω dongles and three that we've had dvi doctors, that we'd have seen at least one consistent or intermittent failure. I propose we deploy 75Ω dongles to all talos-r4 machines because they seem to fix the issues and are easier to deploy than dvi doctors. If we want to have another meeting to discuss this, I can organize one for later this week or early next, otherwise I'd like to move forward with the 75Ω dongles.
That's great news! I'll file a bug to get started on that deployment. I want to do a bit more experimentation in the interim in bug 713746, so I'll leave this open. Bug 714954 covers the 75Ω deployment.
No longer depends on: 714954
Blocks: 714954
No longer depends on: 714954
As noted in bug 713746#c15, a bug in slavealloc has caused the conclusion regarding 75Ω dongles in comment 88 to be based on insufficient data. Because of this, I am closing bug 714954 as INVALID. We can file a new bug if we decide to roll out the 75Ω dongles. I am going to re-enable the 75Ω dongle machines in slavealloc. Aside from snow-018, all the dvi doctor machines were enabled in slavealloc, even though their comments were overwritten with incorrect data.
Remember I'm using snow-014. It won't start buildbot in its current state anyway, but it should still be disabled and commented as such.
Bug 713746 is mostly at an end now, I think. Evidence suggests that a script similar to that given at the end of the bug will do the trick in production -- if it can't set the resolution, it cold-boots the box. As John said in comment 46, what we need is a way to work around the problem, since we've not been able find any way to actually solve it. John, what do you think - should we put something like this into practice and see what it misses?
Attachment #583255 - Flags: review?(coop) → review+
(In reply to Dustin J. Mitchell [:dustin] from comment #92) > John, what do you think - should we put something like this into practice > and see what it misses? There are a bunch of different patches in flight that address parts of this issue: * Attachment #583255 [details] [diff] (in this bug) * Attachment #583849 [details] [diff] (bug 712750) * the script in https://bugzilla.mozilla.org/show_bug.cgi?id=713746#c18 Do we need all these moving parts, or do some patches supercede others? Can we amalgamate some/all of these bugs, please?
(In reply to Chris Cooper [:coop] from comment #93) > (In reply to Dustin J. Mitchell [:dustin] from comment #92) > > John, what do you think - should we put something like this into practice > > and see what it misses? > > There are a bunch of different patches in flight that address parts of this > issue: > > * Attachment #583255 [details] [diff] (in this bug) This is a script that lets us do the cold reboot. Its purpose is to prevent pool size collapse, but does nothing to prevent failure. > * Attachment #583849 [details] [diff] (bug 712750) Extra reporting to help figure out what is going on. > * the script in https://bugzilla.mozilla.org/show_bug.cgi?id=713746#c18 I assume this is a replacement for pmset-reboot.sh mentioned in patch 583255 (puppet changes attached to this bug) > Do we need all these moving parts, or do some patches supercede others? Can > we amalgamate some/all of these bugs, please? The two r+'d patches are what we need if we continue trying to work around instead of solve.
I would like to install DVI Doctors on all rev4 slaves. These devices have been running in production without issue since Dec 15, nearly a month. Not a single detected intermittent failure and none of these machines have been in the consistent failure state since the dvi doctors were installed. I think spending the time and money on a known solution is a lot better than continuing to chase down the problem when after two months, we still don't know what the root problem is or even what triggers it.
Before we do that, let's run them through the testing regime in bug 713746.
(In reply to Dustin J. Mitchell [:dustin] from comment #96) > Before we do that, let's run them through the testing regime in bug 713746. I have disabled the following slaves in slavealloc for your tests. talos-r4-snow-011 talos-r4-snow-018 talos-r4-snow-023 When do you expect results from this test?
I think we've wrapped up startup resolution failures, which is what bug 713746 is about. I'd be happy with simply implementing this cold-boot-on-failure technique. I think that the cost is small enough that installing 75Ω dongles on all rev4's is acceptable too, without further research -- although as I stated above I'm not convinced it's necessary. I don't think we have any good reason to deploy DVI Doctors.
(In reply to Dustin J. Mitchell [:dustin] from comment #98) > I think we've wrapped up startup resolution failures, which is what bug > 713746 is about. I'd be happy with simply implementing this > cold-boot-on-failure technique. I think that the cost is small enough that > installing 75Ω dongles on all rev4's is acceptable too, without further > research -- although as I stated above I'm not convinced it's necessary. I > don't think we have any good reason to deploy DVI Doctors. So I keep going back-and-forth on this depending on whether I've talked to server ops or jhford last. DVI Doctors are known to work but are not cheap, and may entail further work reshuffling minis due to lack of space within racks. If we *need* to do this, I have no problem pulling the trigger on the purchase. However we have software fixes in-hand that (according to https://bugzilla.mozilla.org/show_bug.cgi?id=713746#c27) will work around the known failure cases. Let's go ahead and deploy the software fixes before we incur the hardware costs.
(In reply to Chris Cooper [:coop] from comment #99) > (In reply to Dustin J. Mitchell [:dustin] from comment #98) > > I think we've wrapped up startup resolution failures, which is what bug > > 713746 is about. I'd be happy with simply implementing this > > cold-boot-on-failure technique. I think that the cost is small enough that > > installing 75Ω dongles on all rev4's is acceptable too, without further > > research -- although as I stated above I'm not convinced it's necessary. I > > don't think we have any good reason to deploy DVI Doctors. > > So I keep going back-and-forth on this depending on whether I've talked to > server ops or jhford last. > > DVI Doctors are known to work but are not cheap, and may entail further work > reshuffling minis due to lack of space within racks. If we *need* to do > this, I have no problem pulling the trigger on the purchase. > > However we have software fixes in-hand that (according to > https://bugzilla.mozilla.org/show_bug.cgi?id=713746#c27) will work around > the known failure cases. Let's go ahead and deploy the software fixes before > we incur the hardware costs. This is not a fix for the dongle issues. This is a utility that cleans up after a failure so we don't lose the entire slave pool. The only thing that fixes this issue is the DVI Dongle. The utility from 713746 is not a fix, it does not prevent failure and still allows resolution issues to cause developer impacting test failures.
You'll recall, from that meeting, that we agreed to focus on the startup problems, both in hopes we could fix it, and that we could learn something about the behavior of the minis. We can argue the semantics of "fix", but the point is that the script described in bug 713746 makes the startup problem go away, with any sort of dongle. We've also learned some things about the resolution behavior -- in particular, that it can change without warning. Nothing in those tests indicates that the change can only occur on (warm) reboot, so it's certainly reasonable to think that the failures occur any time. What data do we have to support that? How significant is the evidence that the DVI Doctors fix runtime failures? I only know of the three instances of runtime failures from November and early December cited in https://etherpad.mozilla.org/2011-12-21-dongles-mtng have there been more? I'd like to know how many tests have been run on each kind of dongle over a given time period, and how many runtime failures have occurred for each kind of dongle. We planned, in the meeting, to begin gathering more comprehensive data. How is that going?
During a meeting today we decided to go ahead with DVI Doctors. I have filed bug 720006 to request that we install dvi doctors on all the machines. Marking this bug as resolved because we now know what the solution to these problems are. Please track future work in bug 720006.
Status: NEW → RESOLVED
Closed: 14 years ago
Resolution: --- → FIXED
No longer blocks: 696453
No longer blocks: 696417
No longer blocks: 695679
No longer blocks: 693918
My apologies for being out of the loop here. Where can I see the data that was collected here: (In reply to Dustin J. Mitchell [:dustin] from comment #101) > We planned, in the meeting, to begin gathering more comprehensive data. How > is that going? That lead to this?: (In reply to John Ford [:jhford] from comment #102) > During a meeting today we decided to go ahead with DVI Doctors. I have > filed bug 720006 to request that we install dvi doctors on all the machines.
There was no additional data, which is a shame. However, the combination of * no runtime failures on DVI doctors in production * several (John didn't have a number) failures on both 75Ω and 100Ω dongles * DVI Doctors' good behavior in my tests of the startup failures, and * no discernable difference in startup failure behavior between 75Ω and 100Ω suggests (say, p<0.15) that DVI Doctors *do* fix the problems (not surprisingly), whereas (more surprisingly) dongles do not. The software fix I created for startup failures cannot address runtime failures. While I'm unhappy with the lack of meaningful data on runtime failures here, I don't think that additional (experimental) data will appreciably alter the chosen remedy.
Component: Server Operations: RelEng → RelOps
Product: mozilla.org → Infrastructure & Operations
You need to log in before you can comment on or make changes to this bug.

Attachment

General

Created:
Updated:
Size: