Closed
Bug 700672
(r4-dongles)
Opened 14 years ago
Closed 14 years ago
figure out what's going on with talos-r4 dongles
Categories
(Infrastructure & Operations :: RelOps: General, task)
Tracking
(Not tracked)
RESOLVED
FIXED
People
(Reporter: jhford, Assigned: jhford)
References
()
Details
Attachments
(3 files)
|
693 bytes,
patch
|
coop
:
review+
|
Details | Diff | Splinter Review |
|
781 bytes,
patch
|
mozilla
:
review+
|
Details | Diff | Splinter Review |
|
817 bytes,
patch
|
coop
:
review+
|
Details | Diff | Splinter Review |
We are having lots of issues with dongles on talos-r4-snow machines (haven't been looking for them in talos-r4-lion machines yet).
There are at least two failure cases:
1) machines boot, but resolution is set to 800x600x32 after puppet runs but before test harnesses run
2) machine doesn't think it has a dongle, fails to puppet.
(i verified that these machines don't have 1600x1200 available)
talos-r4-snow-015.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-029.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-030.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-031.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-032.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-034.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-038.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-039.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-040.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-042.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-046.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-050.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-053.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-054.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-058.build.mozilla.org Display 0: 1280x1024x32
I think we need to figure out why this is happening and fix it post-haste.
Some things I think we should try:
A) putting 600Ω resistor across hsync and vsync [1]
B) putting 75Ω resistors across the RGB pins instead of 100Ω [1]
C) use a DVI dongle [2]
[1] http://www.mp3car.com/vbulletin/general-hardware-discussion/1763-how-does-a-video-card-detect-vga-monitor-is-present.html#post11453
[2] http://www.monoprice.com/products/product.asp?c_id=101&cp_id=10110&cs_id=1011003&p_id=3048&seq=1&format=2
| Assignee | ||
Comment 1•14 years ago
|
||
Anyone know if something happened to the slaves listed in comment 1? Was any work done near these machines?
Updated•14 years ago
|
Assignee: nobody → server-ops-releng
Component: Release Engineering → Server Operations: RelEng
QA Contact: release → zandr
Comment 2•14 years ago
|
||
Asked jhford where he was tracking failure frequency, etc, and he didn't have a source other than this bug at present. We *should* start tracking that data though.
jhford: can we extract some loose approximation of the failures/frequency from the rev4 bugs and comments that philor has filed?
Comment 3•14 years ago
|
||
That should be https://bugzilla.mozilla.org/buglist.cgi?quicksearch=su%3Arev4 sw%3Aorange
Comment 4•14 years ago
|
||
No of us have been in that datacenter this week, so I don't think anyone bumped into them. Do you have a list of which hosts fall into which category?
Having the right resolution before puppet seems to indicate a software issue. The machine thinking it doesn't have a dongle at all sounds more like it may be a hardware issue, though. It sounds like there may be multiple issues?
Getting a list of the error conditions of each host (and date stamped each time it occurs) would be very helpful as would knowing if these hosts have ever worked, worked once, etc. The more data we have, the easier it will hopefully be to find patterns.
| Assignee | ||
Comment 5•14 years ago
|
||
I have gone through all the bugs that I can find. Of interest is that twice, machines that were intermittently failing eventually started persistently failing
The following machines have issues:
011,015,018,029,030,031,032,034,038,039,040,046,052,053,054,058,061,069,074 - have failed in the persistent failure state at least once (some were fixed by reseating dongles)
023,033,070,078,080 - have failed in the intermittent state at least once
042,050 - have failed both in the intermittent state *and* persistent state
057,064 - have failed at least twice in the intermittent state
I think we need to look closer at the dongles. Apparently, we are using 100 Ohm resistors across the RGB signals. I wonder if the new motherboard's sensor circuit thinks 100 Ohm is right on the threshold of having no monitor present.
Lets try using a different dongle type as well as a DVI monitor simulator.
The new dongle type has the following pin out:
1 & 6 connected by 75 Ohm resistor
2 & 7 connected by 75 Ohm resistor
3 & 8 connected by 75 Ohm resistor
10 & 13 connected by 600 Ohm resistor
10 & 14 connected by 600 Ohm resistor
The DVI monitor simulator is http://www.monoprice.com/products/product.asp?c_id=101&cp_id=10110&cs_id=1011003&p_id=3048&seq=1&format=2
Please install the new dongle on 042,057,023,011,015
Please install the dvi simulator on 050,064,033,018,029 using an HDMI->DVI converter and program the dvi simulators using a Dell Ultrasharp U2410
raw data: https://docs.google.com/spreadsheet/ccc?key=0AlguvtTDx79WdFZVVVk3RFJ3TzNDS0d1aHN1bE5Tdmc
Severity: normal → major
Updated•14 years ago
|
Assignee: server-ops-releng → arich
Comment 6•14 years ago
|
||
I'll work on this - I'll get the monoprice item delivered to scl1.
I'll order the DVI simulator from monoprice, and a dozen more solder cups - http://search.digikey.com/scripts/DkSearch/dksus.dll?Detail&name=17EHD-015-P-AA-0-00-ND
Jake, how are your soldering skills?
| Assignee | ||
Comment 7•14 years ago
|
||
do we already have the 75Ω and 600Ω resistors?
Comment 8•14 years ago
|
||
I know I am going to regret admitting this but my solder skills are excellent.
Comment 9•14 years ago
|
||
All of these are now showing up with 1600x1200. I think it would be a good idea to check the screen resolution on all of the r4 machines at this point, because there was a lot of crawling around and trying to reseat dongles and push in cables. It's very *very* tight in these racks, so I want to make sure we haven't dislodged anything else.
Some of the ones below were touching metal chassis or other metal bits of the rack, so that could easily account for a short that would temporarily take out the dongle. This isn't the case for all of them, though. Since we keep getting asked to fish around in back of these machines, my concern is that we're fixing some and then breaking others in the process (all it takes is bumping two of them together, or bumping one into a rack, I bet) since they're really fragile.
| Assignee | ||
Comment 10•14 years ago
|
||
We now have new machines with dongle problems:
talos-r4-snow-062.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-065.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-067.build.mozilla.org Display 0: 800x600x32
talos-r4-snow-070.build.mozilla.org Display 0: 1280x1024x32
| Assignee | ||
Comment 11•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #10)
> We now have new machines with dongle problems:
to be clear, these are the only machines not currently at the right resolution
| Assignee | ||
Comment 12•14 years ago
|
||
Almost forgot to check the lion machines. I think this issue is affecting them too. I'd like to concentrate on the snow leopard machines. As this is looking like a hardware issue, I think that we should fix we do for the snow machines should be done for lion as well.
The following lion machines are currently having dongle issues:
talos-r4-lion-049.build.mozilla.org Display 0: 1280x1024x32
talos-r4-lion-051.build.mozilla.org Display 0: 1280x1024x32
talos-r4-lion-062.build.mozilla.org Display 0: 1280x1024x32
Summary: figure out what's going on with talos-r4-snow dongles → figure out what's going on with talos-r4 dongles
Comment 13•14 years ago
|
||
I can't find 600W resistors that have a reasonable tolerance. Here's what I'm looking at:
1 12 17EHD-015-P-AA-0-00-ND CONN HD DB15 MALE GOLD FLASH 0 1.89000 $22.68
2 50 75ADCT-ND RES 75 OHM 1/4W 0.1% MF AXL 0 0.26460 $13.23
3 50 604XBK-ND RES 604 OHM 1/4W 1% METAL FILM 0 0.09100 $4.55
I'm assuming that the 75 ohm resistors are probably the tighter tolerance, and the 600's are pull-up/pull-down's, so 604 += 6.04 (1%) is OK. I haven't ordered these yet, though, so please speak up if this is incorrect.
Comment 14•14 years ago
|
||
I'll reseat these tomorrow:
talos-r4-lion-049.build.mozilla.org Display 0: 1280x1024x32
talos-r4-lion-051.build.mozilla.org Display 0: 1280x1024x32
talos-r4-lion-062.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-062.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-065.build.mozilla.org Display 0: 1280x1024x32
talos-r4-snow-067.build.mozilla.org Display 0: 800x600x32
talos-r4-snow-070.build.mozilla.org Display 0: 1280x1024x32
Updated•14 years ago
|
colo-trip: --- → scl1
Comment 15•14 years ago
|
||
one example calling for 75 ohm resistors:
http://bitclockers.com/forums/index.php?topic=7.0
Not sure where the 600 ohm's figure in.
colo-trip: scl1 → ---
Comment 16•14 years ago
|
||
The DVI Doctor (monoprice) and soldering stuff (comment 13) are ordered, for delivery to 650 castro attn matt larrain.
Comment 17•14 years ago
|
||
(In reply to Dustin J. Mitchell [:dustin] from comment #13)
> I can't find 600W resistors that have a reasonable tolerance.
1% is a reasonable tolerance. 0.1% is crazy exotic precision stuff, which is why they're crazy expensive. (26 cents for a resistor is crazy expensive)
> I'm assuming that the 75 ohm resistors are probably the tighter tolerance,
> and the 600's are pull-up/pull-down's, so 604 += 6.04 (1%) is OK. I haven't
> ordered these yet, though, so please speak up if this is incorrect.
All these are doing is simulating the impedance of the monitor inputs. Nothing we're doing is emulating DDC/EDID, so at best the macs will think there's "something" out there, but won't know what it is.
The RGB video signals expect a 75ohm impedance. I can't find a solid reference for the input impedance for H-sync and V-sync, but "I read it on the internet" class references say ~2k. 600 ohms is rather a lot of load for a driver expecting 2k. Further, I'd be very surprised if anyone designed a system to look for load on such a high-impedance input. High impedance means there's very little load on the driver, which makes detecting said load very difficult.
Add in impedance mismatches due to the various connectors (VGA connectors are ~100ohm impedance, can't even begin to speculate about MiniDP) and this becomes a pretty fuzzy way to detect a monitor.
tl;dr going down to 50-75ohms on the RGB lines may help. I don't expect the resistors on Hsync/Vsync to matter. DVIdoctors will work, but are expensive and annoying to configure. At least they don't need power except for programming.
Comment 18•14 years ago
|
||
Great info, thanks.
Yeah, $0.26 each is a *lot*, but I figured if 100 ohms is out of range, then the tolerances are fairly tight, so 5% would still leave room for ambiguity.
Jake mentioned he's used 68Ω before. From this info, it sounds like the things to try are:
* replace 100-ohm with 75-ohm dongles on five minis, bake for a few days, and
* add the DVI Doctor to one node
if the DVI Doctor doesn't work, we should look for software problems.
if the DVI Doctor works, but the 75-ohm do not,
* add 600Ω resistors to h-sync/v-sync on those cups, and
* try the same with 68-ohm resistors, with and without h-sync/v-sync.
I'd be happy to start with 2kΩ instead, although sources I see around the 'net are using 470Ω and 1kΩ, so maybe that's flexible. Let's hope it doesn't get there.
Jake: solder cups with different resistances will be easily mixed up, so if you can find some easy way to label them, or at least make them quickly visually distinguishable, that'd be great.
Comment 19•14 years ago
|
||
digi-key order is
http://www.fedex.com/Tracking?language=english&cntry_code=us&tracknumbers=0126%200607%201554%20715
(not shipped yet, but was shipped ground so ~end of the week)
Comment 20•14 years ago
|
||
These were grouped pretty physically close to each other:
talos-r4-lion-049
1280x1024x16 1280x1024x32 1600x1200x8 1600x1200x16
talos-r4-lion-051
1280x1024x16 1280x1024x32 1600x1200x8 1600x1200x16
talos-r4-lion-062
1280x1024x16 1280x1024x32 1600x1200x8 1600x1200x16
talos-r4-snow-062
1600x1200x8 1600x1200x16 1600x1200x32 1600x1200x8
talos-r4-snow-065
1600x1200x8 1600x1200x16 1600x1200x32 1600x1200x8
talos-r4-snow-067
1600x1200x8 1600x1200x16 1600x1200x32 1600x1200x8
talos-r4-snow-070
1600x1200x8 1600x1200x16 1600x1200x32 1600x1200x8
We wound up putting in some repaired hardware, so it would be interesting to know if others failed near where we were working.
| Assignee | ||
Comment 21•14 years ago
|
||
(In reply to Amy Rich [:arich] [:arr] from comment #20)
> These were grouped pretty physically close to each other:
They all currently show as being at 1600x1200x32, are these machines that had failures in the spreadsheet?
Comment 22•14 years ago
|
||
Those were the ones that you said needed to be reseated in this bug, so I did so. I just did a grep for 1600x1200|head -1.
Comment 23•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #5)
> I have gone through all the bugs that I can find. Of interest is that
> twice, machines that were intermittently failing eventually started
> persistently failing
> raw data:
> https://docs.google.com/spreadsheet/
> ccc?key=0AlguvtTDx79WdFZVVVk3RFJ3TzNDS0d1aHN1bE5Tdmc
Is there a key that says what means what for this data?
I didn't see any more entries after yesterday, so does that mean that no hosts are currently showing errors?
I'm presuming that puppet is notifying you when the resolution is incorrect, right, so you should know fairly quickly when the permanent failure state happens and a reboot won't fix it.
| Assignee | ||
Comment 24•14 years ago
|
||
(In reply to Amy Rich [:arich] [:arr] from comment #22)
> Those were the ones that you said needed to be reseated in this bug, so I
> did so. I just did a grep for 1600x1200|head -1.
ahhh, that makes sense :)
(In reply to Amy Rich [:arich] [:arr] from comment #23)
> (In reply to John Ford [:jhford] from comment #5)
> > I have gone through all the bugs that I can find. Of interest is that
> > twice, machines that were intermittently failing eventually started
> > persistently failing
>
> > raw data:
> > https://docs.google.com/spreadsheet/
> > ccc?key=0AlguvtTDx79WdFZVVVk3RFJ3TzNDS0d1aHN1bE5Tdmc
>
> Is there a key that says what means what for this data?
its not a well organized spreadsheet, but its a list of machines. each time there is a failure, the bug number is put in the next available column to the right. For machines that don't get stuck in a failing state, a date/time stamp is added as well as the bug number
> I didn't see any more entries after yesterday, so does that mean that no
> hosts are currently showing errors?
it means that I haven't updated the spreadsheet
> I'm presuming that puppet is notifying you when the resolution is incorrect,
> right, so you should know fairly quickly when the permanent failure state
> happens and a reboot won't fix it.
puppet notifying us? we don't have any of that, sadly.
Comment 25•14 years ago
|
||
Going to hand this off to jake since I'm leaving town tonight and he'll be doing the soldering work.
Assignee: arich → jwatkins
Comment 26•14 years ago
|
||
I am recommending we seal all the dongles with an Acrylic Conformal Coating. I've used this stuff before to protect pcb prototypes and it works great. It has a fast cure time that doesn't require heat and its viscosity can be adjusted.
http://www.mgchemicals.com/products/419b.html
http://www.mgchemicals.com/downloads/pdf/specsheets/419l.pdf
Comment 27•14 years ago
|
||
An alternative option to the acrylic: install housing on each connector.
http://www.allelectronics.com/make-a-store/item/DB-15H/DB-15-HOOD/1.html
Comment 28•14 years ago
|
||
Summarizing, we have three things going on here:
* Jake is working on ways of insulating the dongles
* Resistors and whatnot (comment 13) have arrived in mtv1 - on Matt's desk
* DVI Doctor has arrived on 11/16 - on Matt's desk
Jake's got the helm as far as the insulation goes. For the experimental new dongles, the next steps are outlined in comment #18, and subject to revision as Jake sees fit.
colo-trip: --- → scl1
| Assignee | ||
Comment 29•14 years ago
|
||
(In reply to Dustin J. Mitchell [:dustin] from comment #28)
> Summarizing, we have three things going on here:
>
> * Jake is working on ways of insulating the dongles
> * Resistors and whatnot (comment 13) have arrived in mtv1 - on Matt's desk
> * DVI Doctor has arrived on 11/16 - on Matt's desk
>
> Jake's got the helm as far as the insulation goes. For the experimental new
> dongles, the next steps are outlined in comment #18, and subject to revision
> as Jake sees fit.
How are things progressing? Please let me know when these tests are ready so the machines can be taken out of production for installation and testing.
Comment 30•14 years ago
|
||
John, we'll be ready for these on arriving at scl1 tomorrow morning. Can you disable and post a list of slaves that we should put these on? Jake has 5 dongles, plus the DVI Doctor, so we'll need six hosts.
Assignee: jwatkins → jhford
| Assignee | ||
Comment 31•14 years ago
|
||
I guess we aren't going to do 5 simulators as requested in comment 5
Please put the simulator on 042 and a new dongle on 050, 064,033,018,029. I've disabled these slaves in slavealloc.
Comment 32•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #31)
> I guess we aren't going to do 5 simulators as requested in comment 5
The discussion progressed significantly from there, given a lot of new information, and we're working off the more recent plan in comment 18. I didn't get any feedback to the contrary.
Are the numbers in comment 31 snow's or lions?
| Assignee | ||
Comment 33•14 years ago
|
||
comment 31 is referring to snow leopard machines
Comment 34•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #31)
> I guess we aren't going to do 5 simulators as requested in comment 5
>
> Please put the simulator on 042 and a new dongle on 050, 064,033,018,029.
> I've disabled these slaves in slavealloc.
I have installed the dvi Dr on snow-042 and the new 75ohm dongles on snow-050, 064,033,018 and 029.
| Assignee | ||
Comment 35•14 years ago
|
||
All machines in comment 34 have been re-enabled in slavealloc.
Comment 36•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #35)
> All machines in comment 34 have been re-enabled in slavealloc.
Any update on these machines and associated failures, John?
| Assignee | ||
Comment 37•14 years ago
|
||
i am letting these run in production for a while to build up data.
| Assignee | ||
Comment 38•14 years ago
|
||
Well, it looks like we are seeing something odd happening in bug 709436 to talos-r4-snow-064. The failure mode is changing, where the machines are reporting 1600x1200x32 but are failing with symptoms of being at 800x600x32. This is not encouraging for the new dongle design as implemented.
Lets go back to the original plan of having multiple dvi-doctors. Can we overnight 9 additional dvi-doctors to SCL1 to be installed on 057,023,011,015,050,064,033,018,029?
Comment 39•14 years ago
|
||
From what I can tell, this sounds like a software error - either in 'screenresolution', or in the tests themselves. If those are giving different results, then one of them is incorrect, and we should stop trusting that result.
Also, it doesn't look like https://docs.google.com/spreadsheet/ccc?key=0AlguvtTDx79WdFZVVVk3RFJ3TzNDS0d1aHN1bE5Tdmc#gid=0 has been kept up to date, which makes it hard to see what patterns there might be here.
DVI Doctors are useful for experimentation and data-gathering, but will not work well in production - they are large and heavy and will most likely fall off of the connectors or pull them out of the minis.
We can get more DVI Doctors on Monday, but given the above (likely software problem, incomplete data, and not a production solution), can you give some details on how that will help to figure out what's wrong here?
Comment 40•14 years ago
|
||
DVI Doctors are ordered:
Order Number: 5491824 - Your Order Detail
The courier price ("Overnight Express") was cheapest, so I went with that option. They don't ship on weekends, so that means it will be over-Monday-night, and arrive in mtv1 on Tuesday.
Jake, let's plan to get these programmed in mtv1 on Tuesday afternoon, after they arrive.
Comment 41•14 years ago
|
||
(In reply to Dustin J. Mitchell [:dustin] from comment #39)
> From what I can tell, this sounds like a software error - either in
> 'screenresolution', or in the tests themselves. If those are giving
> different results, then one of them is incorrect, and we should stop
> trusting that result.
Is that the only possible explanation? Since my Mac changes resolution multiple times per day as I plug in and unplug the external monitor, I would have said the most likely explanation for screenresolution saying one thing at the start of every run, and then sometimes it being the same and sometimes it being different after n minutes when we hit mochitest-chrome, and sometimes it being the same as it was for screenresolution after m minutes in mochitest-browser-chrome and other times it being the same as it was for mochitest-chrome and yet other times being different than it was for either screenresolution or for mochitest-chrome, would be that the resolution is changing during the course of the run.
Comment 42•14 years ago
|
||
It's by no means the only explanation. I'm hoping we can get together and share data on what's gone wrong, and what possible solutions we have.
Comment 43•14 years ago
|
||
If we're going to gather data about how the dongles are working, we probably need some way of keeping those slaves working. If buildapi/recent/ is telling the truth about last jobs,
https://build.mozilla.org/buildapi/recent/talos-r4-snow-018 - 2011-12-07 22:44
https://build.mozilla.org/buildapi/recent/talos-r4-snow-029 - current
https://build.mozilla.org/buildapi/recent/talos-r4-snow-033 - 2011-12-07 23:04
https://build.mozilla.org/buildapi/recent/talos-r4-snow-042 - current (after a two day layoff ending last night)
https://build.mozilla.org/buildapi/recent/talos-r4-snow-050 - 2011-12-04 09:49
https://build.mozilla.org/buildapi/recent/talos-r4-snow-064 - current (though it's blowing so many jobs it probably shouldn't be, bug 709436)
Comment 44•14 years ago
|
||
Good point. If the talos-r4-snow pool is so under-utilized that the slaves sit idle for days, maybe we should disable half (or some other fraction) of them temporarily?
Comment 45•14 years ago
|
||
It was late, but I think jhford said that 042 had lost its connection to its master after a reconfig, and he rebooted it - if I heard him right, 018 and 033 are probably in the same state (and maybe even 050 from a previous incident).
| Assignee | ||
Comment 46•14 years ago
|
||
(In reply to Dustin J. Mitchell [:dustin] from comment #39)
> From what I can tell, this sounds like a software error - either in
> 'screenresolution', or in the tests themselves. If those are giving
> different results, then one of them is incorrect, and we should stop
> trusting that result.
>
> Also, it doesn't look like
> https://docs.google.com/spreadsheet/
> ccc?key=0AlguvtTDx79WdFZVVVk3RFJ3TzNDS0d1aHN1bE5Tdmc#gid=0 has been kept up
> to date, which makes it hard to see what patterns there might be here.
I put that together not as the authoritative data source, but as something for me to work with to get numbers.
> DVI Doctors are useful for experimentation and data-gathering, but will not
> work well in production - they are large and heavy and will most likely fall
> off of the connectors or pull them out of the minis.
I think it is too early to say that they will not work in production. DVI Doctors can use the HDMI->DVI dongle that is already included in R4 and R5 minis. We have shown that the minidp connectors do not reliably make a good connection. The HDMI cables are significantly more solid.
There is shelf space behind each mini, I think we can find a way to mount this box to each shelf.
If DVI Doctors present a clear solution to this problem, I think we ought to consider them the long term fix. I don't want to spend yet more time finding the cheapest solution if we have a solid solution.
> We can get more DVI Doctors on Monday, but given the above (likely software
> problem, incomplete data, and not a production solution), can you give some
> details on how that will help to figure out what's wrong here?
software vs. hardware is not an important distinction. What is import to differentiate is whether we have control to fix the defect. If the bug is in the display apis, driver, OS or hardware, there is nothing we can do to fix it and must work around it.
I don't think we can call the dvi doctor "not a production solution" yet. It sounds a lot more like a 'production' solution than a minidp plug that consistently does not make a solid connection.
(In reply to Phil Ringnalda (:philor) from comment #45)
> It was late, but I think jhford said that 042 had lost its connection to its
> master after a reconfig, and he rebooted it - if I heard him right, 018 and
> 033 are probably in the same state (and maybe even 050 from a previous
> incident).
Yes, 042 had lost its connection to the master, probably a network 'blip'. It has been back in the pool and has not experienced any dongle related issues so far, though, its too early to tell if the dvi doctor fixes the problem.
(In reply to Dustin J. Mitchell [:dustin] from comment #44)
> Good point. If the talos-r4-snow pool is so under-utilized that the slaves
> sit idle for days, maybe we should disable half (or some other fraction) of
> them temporarily?
More likely, if they aren't reaching production masters, they are stuck in the failure state of having the wrong resolution.
Comment 47•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #46)
> If DVI Doctors present a clear solution to this problem, I think we ought to
> consider them the long term fix. I don't want to spend yet more time
> finding the cheapest solution if we have a solid solution.
That may be fair - let's see how the next 10 work out.
> More likely, if they aren't reaching production masters, they are stuck in
> the failure state of having the wrong resolution.
Correct -- bug 709596. At this point, that bug is blocking gathering additional data on this problem, because any host that comes up with an incorrect resolution hangs there until manually fixed.
| Assignee | ||
Comment 48•14 years ago
|
||
I have a new version of screenresolution which prints more information. The logs are visible on stderr as well as the standard apple logging mechanism. Each run has at least the argv printed to logs, so we can see exactly how many times the program is being invoked and when.
~/software/screenresolution $ ./screenresolution
2011-12-13 11:27:44.945 screenresolution[19469:707] starting screenresolution argv=./screenresolution
2011-12-13 11:27:44.946 screenresolution[19469:707] Incorrect command line
~/software/screenresolution $ ./screenresolution get
2011-12-13 11:27:49.953 screenresolution[19471:707] starting screenresolution argv=./screenresolution get
2011-12-13 11:27:49.958 screenresolution[19471:707] Display 0: 1920x1200x32
~/software/screenresolution $ ./screenresolution set
2011-12-13 11:27:52.305 screenresolution[19472:707] starting screenresolution argv=./screenresolution set
~/software/screenresolution $ ./screenresolution set 1600x1200x32
2011-12-13 11:27:57.833 screenresolution[19473:707] starting screenresolution argv=./screenresolution set 1600x1200x32
2011-12-13 11:27:57.837 screenresolution[19473:707] set mode on display 0 to 1600x1200x32
~/software/screenresolution $ ./screenresolution list
2011-12-13 11:28:06.904 screenresolution[19476:707] starting screenresolution argv=./screenresolution list
Available Modes on Display 0
1920x1200x16 1920x1200x32 1920x1200x30 960x600x16
960x600x32 960x600x30 1680x1050x16 1680x1050x32
1680x1050x30 1600x1200x16 1600x1200x32 1600x1200x30
1600x1200x16 1600x1200x32 1600x1200x30 1280x1024x16
1280x1024x32 1280x1024x30 1280x1024x16 1280x1024x32
1280x1024x30 1152x720x16 1152x720x32 1152x720x30
1024x768x16 1024x768x32 1024x768x30 1024x768x16
1024x768x32 1024x768x30 1024x640x16 1024x640x32
1024x640x30 1280x800x16 1280x800x32 1280x800x30
800x600x16 800x600x32 800x600x30 800x600x16
800x600x32 800x600x30 800x500x16 800x500x32
800x500x30 640x480x16 640x480x32 640x480x30
640x480x16 640x480x32 640x480x30 720x480x16
720x480x32 720x480x30 720x480x16 720x480x32
720x480x30 1280x960x16 1280x960x32 1280x960x30
1280x960x16 1280x960x32 1280x960x30 1344x1008x16
1344x1008x32 1344x1008x30 1344x840x16 1344x840x32
1344x840x30 1600x1000x16 1600x1000x32 1600x1000x30
Attachment #581351 -
Flags: review?(coop)
Comment 49•14 years ago
|
||
Jake has the dongles and will get them programmed today for installation on Thursday.
| Assignee | ||
Comment 50•14 years ago
|
||
(In reply to Dustin J. Mitchell [:dustin] from comment #49)
> Jake has the dongles and will get them programmed today for installation on
> Thursday.
This doesn't leave us with a week before the meeting on Dec 21. Is it possible to install them today or tomorrow instead?
Comment 51•14 years ago
|
||
Nobody's onsite until Thursday. We'll still have 6 days', which should give us something to guess on -- even if that means we agree to share the data next Thursday and discuss in IRC.
Updated•14 years ago
|
Alias: r4-dongles
Comment 52•14 years ago
|
||
As a point of order, the DVI doctor we currently have installed is connected via a standard DVI cable and a HDMI->DVI adapter. It could also use a MiniDP->DVI adapter, although my own experience with the video on these devices suggests that the two outputs are not symmetrical, so that may change the results.
All of the dongles with solder cups are MiniDP->VGA, so if we choose to go with DVI Doctors, we will need more adapters as well, and probably some shorter DVI cables.
| Assignee | ||
Comment 53•14 years ago
|
||
HDMI->DVI cables were included with every rev4 mini we bought [1]. Lets reuse the ones we already own. My preference is that we use the HDMI->DVI converter because we already have them, because they are official apple parts and because they seem to have a significantly more solid connection than mini-dp cables. We could also try out these: http://www.monoprice.com/products/product.asp?c_id=102&cp_id=10231&cs_id=1023104&p_id=2661&seq=1&format=2 if we decide to go with DVI Doctors.
[1] http://support.apple.com/kb/SP585
Comment 54•14 years ago
|
||
Yeah, Zandr pointed that out right after I posted the comment.
Comment 55•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #46)
> I think it is too early to say that they will not work in production.
It is also too early to say that they will. I'd like to get to a DC and look at how they're installed, but I agree with Dustin that they are mechanically challenging. They are also labor intensive to program, and we don't know anything about their failure modes yet.
> DVI Doctors can use the HDMI->DVI dongle that is already included in R4 and R5
> minis. We have shown that the minidp connectors do not reliably make a good
> connection. The HDMI cables are significantly more solid.
"Shown" implies that you have data. If you have data that shows that the HDMI connection is more reliable than the Mini-DP, then you haven't shared it. In order to control for the connector, we'd have to be using the same solution (DVI Doctor or resistor dongle) on both connectors. As HDMI has no analog, that means we'd need to test with DVI Doctors on Mini-DP to run that experiment.
> There is shelf space behind each mini, I think we can find a way to mount
> this box to each shelf.
I'm glad you think so. I think those shelves are already fairly full of cables, and probably don't make a good mounting location. Someone who has actually installed a DVI Doctor in a full rack should comment here. I'm particularly concerned about servicablility (which is already pretty poor) once we get a number of them packed together.
Remember also that since both the HDMI adapter and the DVI-Doctor have female connectors, we need a short DVI cable between them. This is starting to be a lot of hardware to cram into that shelf, and we'll probably have to engineer some sort of mounting rail inside the rack.
Were the machines that are getting the DVI Doctors tomorrow selected for adjacency? Could we do that, and figure out if we can actually install these rationally on every unit?
> If DVI Doctors present a clear solution to this problem, I think we ought to
> consider them the long term fix. I don't want to spend yet more time
> finding the cheapest solution if we have a solid solution.
I think we need to evaluate data from all of the solutions before we start jumping to conclusions, and weigh that against the operational considerations of each solution. There are costs to each solution that go beyond releng's time and the bill from Monoprice.
> software vs. hardware is not an important distinction.
Uhh, what? If you are declining to identify the problem, I find it very difficult to understand how you expect to come up with a solution.
> I don't think we can call the dvi doctor "not a production solution" yet.
> It sounds a lot more like a 'production' solution than a minidp plug that
> consistently does not make a solid connection.
I don't think we can call *anything* a production solution yet. I do think it's very easy for you to discount the operational concerns because they aren't your problem. As I said, there are costs to each solution that go beyond your time and the cost of the DVI doctors.
> More likely, if they aren't reaching production masters, they are stuck in
> the failure state of having the wrong resolution.
Do we have logging that could determine that (either before or after the improved tool in comment 48)?
There's too much speculation and not enough data here, let's fix that.
Updated•14 years ago
|
Attachment #581351 -
Flags: review?(coop) → review+
Comment 56•14 years ago
|
||
I will need a list of slaves to attach the 9 new DVI Doctors. John, can you provide this please? I will be on site at SCL1 today.
| Assignee | ||
Comment 57•14 years ago
|
||
057,023,011,015,050,064,033,018,029
(that's from comment 38, but there's been a lot of churn since then)
Comment 58•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #57)
> 057,023,011,015,050,064,033,018,029
As I asked in comment #55:
> Were the machines that are getting the DVI Doctors tomorrow selected for
> adjacency? Could we do that, and figure out if we can actually install these
> rationally on every unit?
Yes, no, maybe?
Comment 59•14 years ago
|
||
9 DVI Doctors have been installed on 057,023,011,015,050,064,033,018,029.
| Assignee | ||
Comment 60•14 years ago
|
||
(In reply to Zandr Milewski [:zandr] from comment #55)
> (In reply to John Ford [:jhford] from comment #46)
> > I think it is too early to say that they will not work in production.
>
> It is also too early to say that they will. I'd like to get to a DC and look
> at how they're installed, but I agree with Dustin that they are mechanically
> challenging. They are also labor intensive to program, and we don't know
> anything about their failure modes yet.
Yes, that's why we are trying to gather data here. I wouldn't want to go through the trouble of installing them on 160 machines to find out that they don't work.
> > DVI Doctors can use the HDMI->DVI dongle that is already included in R4 and R5
> > minis. We have shown that the minidp connectors do not reliably make a good
> > connection. The HDMI cables are significantly more solid.
>
> "Shown" implies that you have data. If you have data that shows that the
> HDMI connection is more reliable than the Mini-DP, then you haven't shared
> it. In order to control for the connector, we'd have to be using the same
> solution (DVI Doctor or resistor dongle) on both connectors. As HDMI has no
> analog, that means we'd need to test with DVI Doctors on Mini-DP to run that
> experiment.
Matt and I debugged this with multiple computers. It turns out that even though the dongles were fully inserted, they weren't being picked up by the computer. Matt tried pushing the dongle in to make sure it was fully inserted. The dongle still wasn't being detected by the computer. The only fix was for Matt to completely remove and reinsert the dongle. This affected a bunch of the slaves, multiple times.
To clarify, both my aggregation and the raw data are public.
> > There is shelf space behind each mini, I think we can find a way to mount
> > this box to each shelf.
>
> I'm glad you think so. I think those shelves are already fairly full of
> cables, and probably don't make a good mounting location. Someone who has
> actually installed a DVI Doctor in a full rack should comment here. I'm
> particularly concerned about servicablility (which is already pretty poor)
> once we get a number of them packed together.
>
> Remember also that since both the HDMI adapter and the DVI-Doctor have
> female connectors, we need a short DVI cable between them. This is starting
> to be a lot of hardware to cram into that shelf, and we'll probably have to
> engineer some sort of mounting rail inside the rack.
As I mentioned in comment 53, we could look at HDMI->DVI cables.
> Were the machines that are getting the DVI Doctors tomorrow selected for
> adjacency? Could we do that, and figure out if we can actually install these
> rationally on every unit?
They were selected for having demonstrated the issues we are trying to correct for. I still think this is the correct basis for selection.
> > If DVI Doctors present a clear solution to this problem, I think we ought to
> > consider them the long term fix. I don't want to spend yet more time
> > finding the cheapest solution if we have a solid solution.
>
> I think we need to evaluate data from all of the solutions before we start
> jumping to conclusions, and weigh that against the operational
> considerations of each solution. There are costs to each solution that go
> beyond releng's time and the bill from Monoprice.
Yes, I have been asking to get the DVI Doctors installed since November 9 so we could start collecting data on possible fixes. I am also aware that there are other costs.
> > software vs. hardware is not an important distinction.
>
> Uhh, what? If you are declining to identify the problem, I find it very
> difficult to understand how you expect to come up with a solution.
I believe what I said was:
(comment 46)
> software vs. hardware is not an important distinction. What is import to
> differentiate is whether we have control to fix the defect. If the bug is
> in the display apis, driver, OS or hardware, there is nothing we can do to
> fix it and must work around it.
I don't see anything there that suggests that I am uninterested in figuring out what is causing these problems. In fact, I think finding out what is happening is so important that I filed this very bug to figure out whats going on with the dongles! What I am trying to say is that if the problem turns out to be something outside of our control, we need to work around it.
> > I don't think we can call the dvi doctor "not a production solution" yet.
> > It sounds a lot more like a 'production' solution than a minidp plug that
> > consistently does not make a solid connection.
>
> I don't think we can call *anything* a production solution yet. I do think
> it's very easy for you to discount the operational concerns because they
> aren't your problem. As I said, there are costs to each solution that go
> beyond your time and the cost of the DVI doctors.
I think its pretty obvious that any solution is going to have costs in excess of acquisition. I think we also need to keep in mind the opportunity cost of working on this problem instead of others as well as the cost of wasted developer time dealing with fallout from this issue.
I don't appreciate your characterization and don't think its correct. I have jumped in many times to help out with things that aren't "my problem". I also don't think that this is relevant to the discussion at hand.
> > More likely, if they aren't reaching production masters, they are stuck in
> > the failure state of having the wrong resolution.
>
> Do we have logging that could determine that (either before or after the
> improved tool in comment 48)?
The improvements in comment 48 concern logging in the actual application. We don't have logging to check when the machines get into the failing state.
> There's too much speculation and not enough data here, let's fix that.
Getting data to see what's going wrong is the purpose of this bug. What you are referring to as speculation, I can only assume is people coming up with ideas and trying to implement a test for their idea. If you have a solution, please, do share it!
Comment 61•14 years ago
|
||
I'm curious is all r4 machines are exhibiting these problems, or just a select number (and if those issues are ongoing. John, are you tracking each occurrence now?). If there are machines that are not exhibiting the problems, are we attaching the same hardware solutions we're testing to a control group of functioning machines?
Comment 62•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #60)
> Matt and I debugged this with multiple computers. It turns out that even
> though the dongles were fully inserted, they weren't being picked up by the
> computer. Matt tried pushing the dongle in to make sure it was fully
> inserted. The dongle still wasn't being detected by the computer. The only
> fix was for Matt to completely remove and reinsert the dongle. This
> affected a bunch of the slaves, multiple times.
This might be the most important thing that has been said in this bug, and I think we've all been misinterpreting that information.
To me, this doesn't say anything about connector reliability, but rather indicates that a state change (unplugging and replugging the dongle) is required to resolve the failure.
Beyond that, we have no data on HDMI connector reliability, so we can't make any assertions about the relative reliability of HDMI vs. MiniDP.
> To clarify, both my aggregation and the raw data are public.
In comment 46, you indicated that the aggregation was not authoritative. Is there an authoritative aggregation?
> As I mentioned in comment 53, we could look at HDMI->DVI cables.
Yup, and I haven't found a source for those shorter than 3'. You aren't going to get two 3' cables and two DVI doctors onto a shelf that's already full of several cables.
> They were selected for having demonstrated the issues we are trying to
> correct for. I still think this is the correct basis for selection.
Fair enough.
> I don't appreciate your characterization and don't think its correct. I
> have jumped in many times to help out with things that aren't "my problem".
> I also don't think that this is relevant to the discussion at hand.
It's quite relevant. It is very easy to ignore the costs of any solution that are outside your organization. Every group in this organization is guilty of that to some degree.
But taking a step back and thinking about it, the operational ramifications of DVI Doctors far outweigh some additional debugging effort on the resistor-based solution. We think that we're having trouble with connector reliability, so we're going to add more connectors in a quest to fix it? From an operations standpoint that's the wrong direction.
> > > More likely, if they aren't reaching production masters, they are stuck in
> > > the failure state of having the wrong resolution.
> >
> > Do we have logging that could determine that (either before or after the
> > improved tool in comment 48)?
>
> The improvements in comment 48 concern logging in the actual application.
> We don't have logging to check when the machines get into the failing state.
Do we have any evidence that supports the assertion that machines that aren't talking to the production masters are in that state because of a resolution problem? That's the gap in my understanding.
> Getting data to see what's going wrong is the purpose of this bug. What you
> are referring to as speculation, I can only assume is people coming up with
> ideas and trying to implement a test for their idea. If you have a
> solution, please, do share it!
I've been pointing out experimental design issues. I just want to be very careful that we are supporting the conclusions we draw, which I've not seen in this bug so far.
Comment 63•14 years ago
|
||
(In reply to Zandr Milewski [:zandr] from comment #62)
> (In reply to John Ford [:jhford] from comment #60)
>
> > The only
> > fix was for Matt to completely remove and reinsert the dongle. This
> > affected a bunch of the slaves, multiple times.
>
> This might be the most important thing that has been said in this bug, and I
> think we've all been misinterpreting that information.
>
> To me, this doesn't say anything about connector reliability, but rather
> indicates that a state change (unplugging and replugging the dongle) is
> required to resolve the failure.
Thinking a bit more about this. I believe that this means the connector works *fine*. This raises a couple of questions, and suggests an experiment.
Question: Does this error condition survive a reboot?
If so, we should try the following experiment:
If it isn't already, set "restart after power failure" (systemsetup -setrestartpowerfailure on). IMO, this should be set anyway.
Wait for a machine to enter the 'bad' state, then:
Shut down the machine (not a restart)
Do a PDU reset. Let the machine come back up.
If that fixes the problem, then we know that a warm-reset is not resetting the parts of the gfx controller that are hosed, and a cold-reset does.
I'm not a fan of doing PDU resets since it slows down the reboot cycle, but I'd take it over DVI doctors.
Comment 64•14 years ago
|
||
(In reply to Zandr Milewski [:zandr] from comment #63)
> Shut down the machine (not a restart)
>
> Do a PDU reset. Let the machine come back up.
A little further reading suggests this might not work, that the restart-after-power-cut only works if the system was on when it lost power. Testing this while onsite would be a good thing. :D
| Assignee | ||
Comment 65•14 years ago
|
||
(In reply to Amy Rich [:arich] [:arr] from comment #61)
> I'm curious is all r4 machines are exhibiting these problems, or just a
> select number (and if those issues are ongoing. John, are you tracking each
> occurrence now?). If there are machines that are not exhibiting the
> problems, are we attaching the same hardware solutions we're testing to a
> control group of functioning machines?
There isn't much historic data on machines that don't make it to buildbot right now. Intermittent failures of machines that have made it past the puppet check for resolution are being tracked in bugs 693918, 695679, 696417 and 696453.
I have filed:
bug 711374 for talos-r4-snow-018
bug 711376 for talos-r4-snow-012
bug 711382 for reseating dongles on 42 machines.
In the process for gathering this data, I have found a new failure mode. The quartz window server is failing, but ssh continues to work. In this mode, I can VNC into the machine but the VNC screen is solid blue with a working cursor and nothing else. One machine initially had broken sshd along with the blue screen, but eventually allowed me to log in. This failure mode seems to recover after a reboot initiated over ssh. I saw this mode on machines with working dongles, with malfunctioning dongles and with dvi doctors.
I've created an etherpad to track instances of the consistent failure state with data starting today. Updating it is a mostly manual task (automatically generated data that requires human processing and verification) and takes a while to do.
https://etherpad.mozilla.org/rev4-dongle-log
Comment 66•14 years ago
|
||
bug 711382 looks like it contains a bunch of machines that haven't been seen before. I only have the spreadsheet linked in comment 5, but from that data none of the lion machines and only 6 of 22 snow machines are repeat offenders.
This suggests that the answer to :arr's question in comment 61 is 'no'. This appears to be across the entire pool, and not restricted to certain machines.
None of the machines in bug 711382 have 75ohm dongles, but without historical data on the failures, it's hard to evaluate that.
I also don't understand the dependency on bug 711374 and bug 711376 here. The procedure requested in those bugs will not produce any data about resolution or dongle state.
Comment 67•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #65)
> https://etherpad.mozilla.org/rev4-dongle-log
Is there anything to indicate that modes D and E are in any way related to dongles or screen resolution?
I could see that mode F is the result of resolutions changing and the window server getting confused.
| Assignee | ||
Comment 68•14 years ago
|
||
(In reply to Zandr Milewski [:zandr] from comment #67)
> (In reply to John Ford [:jhford] from comment #65)
> > https://etherpad.mozilla.org/rev4-dongle-log
>
> Is there anything to indicate that modes D and E are in any way related to
> dongles or screen resolution?
not specifically, just logging as much info as I can and these are fairly common cases
> I could see that mode F is the result of resolutions changing and the window
> server getting confused.
Yah, I think that's very reasonable, especially considering that the screens go blue like that during mode changes.
(In reply to Zandr Milewski [:zandr] from comment #66)
> I also don't understand the dependency on bug 711374 and bug 711376 here.
> The procedure requested in those bugs will not produce any data about
> resolution or dongle state.
There may or may not be something wrong with the dongles on these machines. I am happy to unlink the bugs if it turns out that the dongle is not the cause of problems on this machine.
Comment 69•14 years ago
|
||
I just found out that we *replaced* the 75ohm dongles with DVI-Doctors, ending the 75ohm experiment. I don't believe that we have sufficient data to draw any conclusions, therefore you should not take the next sentence seriously, but...
No mini with a 75ohm resistor dongle has failed.
I would reocmmend selecting some repeat offenders and swapping in 75ohm dongles.
Comment 70•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #68)
> There may or may not be something wrong with the dongles on these machines.
> I am happy to unlink the bugs if it turns out that the dongle is not the
> cause of problems on this machine.
Nothing in those bugs says anything about dongles one way or the other. I would recommend unlinking them.
Comment 71•14 years ago
|
||
(In reply to Zandr Milewski [:zandr] from comment #69)
> No mini with a 75ohm resistor dongle has failed.
talos-r4-snow-064?
Comment 72•14 years ago
|
||
(In reply to Phil Ringnalda (:philor) from comment #71)
> talos-r4-snow-064?
Oh, good catch. snow-064 shows as Mode B in the etherpad.
Hm. I do wish we had 75ohm dongles still in service.
Comment 73•14 years ago
|
||
It occurs to me I have some relevant anecdotal evidence from home. I have an r5 mini with two screens, and on wake from sleep in some significant proportion of cases, it will only detect one of the screens. "Detect Displays" will generally fix this. I wonder if there's a way to trigger "Detect Displays" from software? That might provide a path to a solution here.
Comment 74•14 years ago
|
||
Haven't found anything that's part of the distribution, but I did find this: http://code.google.com/p/detectdisplays/
| Assignee | ||
Comment 75•14 years ago
|
||
(In reply to Zandr Milewski [:zandr] from comment #74)
> Haven't found anything that's part of the distribution, but I did find this:
> http://code.google.com/p/detectdisplays/
tried that and it didn't seem to change anything:
talos-r4-lion-012:~ cltbld$ screenresolution list
Available Modes on Display 0
1x1x8 1x1x16 1x1x32 1x1x64
1x1x96 800x600x8 800x600x16 800x600x32
800x600x64 800x600x96 1024x768x8 1024x768x16
1024x768x32 1024x768x64 1024x768x96 1280x1024x8
1280x1024x16 1280x1024x32 1280x1024x64 1280x1024x96
1680x1050x8 1680x1050x16 1680x1050x32 1680x1050x64
1680x1050x96 1280x1024x32
talos-r4-lion-012:~ cltbld$ ./detectdisplays
talos-r4-lion-012:~ cltbld$ screenresolution list
Available Modes on Display 0
1x1x8 1x1x16 1x1x32 1x1x64
1x1x96 800x600x8 800x600x16 800x600x32
800x600x64 800x600x96 1024x768x8 1024x768x16
1024x768x32 1024x768x64 1024x768x96 1280x1024x8
1280x1024x16 1280x1024x32 1280x1024x64 1280x1024x96
1680x1050x8 1680x1050x16 1680x1050x32 1680x1050x64
1680x1050x96 1280x1024x32
| Assignee | ||
Comment 76•14 years ago
|
||
(In reply to Zandr Milewski [:zandr] from comment #72)
> (In reply to Phil Ringnalda (:philor) from comment #71)
>
> > talos-r4-snow-064?
>
> Oh, good catch. snow-064 shows as Mode B in the etherpad.
>
> Hm. I do wish we had 75ohm dongles still in service.
See also: bug 709436. This slave exhibited all sorts of new problems on the 75Ω dongle.
Comment 77•14 years ago
|
||
https://bugzilla.mozilla.org/show_bug.cgi?id=711382#c9 is relevant.
Cold starts recover machines from this state without touching the dongle.
| Assignee | ||
Comment 78•14 years ago
|
||
Changing from an onlyif to unless means that we are explicitly search for a string instead of searching for its absence. This should be generally better. This also redirects the output from NSLog from stderr to stdout for grep to read.
The relevant sections of the puppet type manifest are:
-onlyif: If this parameter is set, then this exec will only run if the command returns 0.
-unless: If this parameter is set, then this exec will run unless the command returns 0
Attachment #582977 -
Flags: review?
| Assignee | ||
Updated•14 years ago
|
Attachment #582977 -
Flags: review? → review?(aki)
Updated•14 years ago
|
Attachment #582977 -
Flags: review?(aki) → review+
| Assignee | ||
Comment 79•14 years ago
|
||
(In reply to Zandr Milewski [:zandr] from comment #77)
> https://bugzilla.mozilla.org/show_bug.cgi?id=711382#c9 is relevant.
>
> Cold starts recover machines from this state without touching the dongle.
Moving discussion over here, as that bug will go away when the machines are operational again.
(In reply to Zandr Milewski [:zandr] from comment #9)
> (In reply to John Ford [:jhford] from comment #8)
>
> > Yes, it recovered but is now broken again. Its last test was at "Fri Dec 16
> > 14:36:44 2011" (pst). This test was successful and at the correct
> > resolution.
>
> Excellent. We now know two more things:
>
> 1) It has nothing to do with connectors on the dongles.
Had the fix stuck, I'd completely agree with you. Given that the machine broke again so soon, I wouldn't personally rule this out as part of the problem. It seems we are right on the edge of what causes the machine not to detect a display.
> 2) Unpleasant though it may be, a cold reboot will recover a machine.
>
> I'll play around with pmset to see if I can come up with a solution that
> will let us do a shutdown instead of a reboot after each test.
Yep, looks like pmset has 'acwake' and 'autorestart' which look like they are both useful. There is also 'shutdown -u' which halts the machine but leaves the power on so the power loss setting properly turns the machine on when the pdu gives it power.
That said, I think we should consider ethernet PDUs + pmset + shutdown -u a last resort if the 75Ω dongles and DVI Doctor don't work. A solution like this adds a lot of complexity to automate it and a lot of ongoing human-time if we keep it manual.
Comment 80•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #79)
> Had the fix stuck, I'd completely agree with you. Given that the machine
> broke again so soon, I wouldn't personally rule this out as part of the
> problem.
I don't see that this has anything to do with the connectors. We recovered 8 out of 8 machines without touching the connector, and we've established previously that wiggling the connectors doesn't help, only removing and reinstalling them helps. That's the same sort of state change we're going through by doing a power cycle.
> It seems we are right on the edge of what causes the machine not
> to detect a display.
This seems to be the case, but the machines we power cycled all had 100Ω dongles, right?
> Yep, looks like pmset has 'acwake' and 'autorestart' which look like they
> are both useful. There is also 'shutdown -u' which halts the machine but
> leaves the power on so the power loss setting properly turns the machine on
> when the pdu gives it power.
I wasn't going that direction, actually. 'acwake' doesn't apply, and as you say, autorestart would require shutdown -u and a power cycle. More in a bit.
> That said, I think we should consider ethernet PDUs + pmset + shutdown -u a
> last resort if the 75Ω dongles and DVI Doctor don't work.
I had been under the impression you'd ruled out 75Ω dongles, glad that's not the case. I really want to rule out DVI Doctors. ;)
> A solution like
> this adds a lot of complexity to automate it and a lot of ongoing human-time
> if we keep it manual.
Using PDU's sure. But there's another way:
pmset schedule wakeorpoweron "`date -v+1M "+%m/%d/%y %H:%M:%S"`" && shutdown -h now
This sets an 'alarm' to power on the machine at now +1 minute and then does a shutdown.
I believe that this is equivalent to the "Test 1" in bug 711382, doing a clean shutdown and then pressing the power button.
| Assignee | ||
Comment 81•14 years ago
|
||
(In reply to Zandr Milewski [:zandr] from comment #80)
> (In reply to John Ford [:jhford] from comment #79)
> > It seems we are right on the edge of what causes the machine not
> > to detect a display.
>
> This seems to be the case, but the machines we power cycled all had 100Ω
> dongles, right?
Yep, they were the 100Ω ones.
> Using PDU's sure. But there's another way:
>
> pmset schedule wakeorpoweron "`date -v+1M "+%m/%d/%y %H:%M:%S"`" && shutdown
> -h now
>
> This sets an 'alarm' to power on the machine at now +1 minute and then does
> a shutdown.
>
> I believe that this is equivalent to the "Test 1" in bug 711382, doing a
> clean shutdown and then pressing the power button.
Can you test this?
That looks neat. What happens if the machine takes longer than 1 minute to shut down? Can we set more than one alarm time? My reading of the manpage suggests we can. If that's the case, we could schedule a wakeorpoweron for 1, 5 and 60 minutes.
Before we move to this, we should have a machine continually reboot by this method to see how often or if it fails.
| Assignee | ||
Comment 82•14 years ago
|
||
talos-r4-snow-013:~ cltbld$ screenresolution list
2011-12-20 07:53:42.407 screenresolution[1325:903] starting screenresolution argv=screenresolution list
Available Modes on Display 0
1x1x8 1x1x16 1x1x32 1x1x64
1x1x96 800x600x8 800x600x16 800x600x32
800x600x64 800x600x96 1024x768x8 1024x768x16
1024x768x32 1024x768x64 1024x768x96 1280x1024x8
1280x1024x16 1280x1024x32 1280x1024x64 1280x1024x96
1680x1050x8 1680x1050x16 1680x1050x32 1680x1050x64
1680x1050x96 1280x1024x32 talos-r4-snow-013:~ cltbld$ screenresolution get
2011-12-20 07:53:51.752 screenresolution[1326:903] starting screenresolution argv=screenresolution get
2011-12-20 07:53:51.759 screenresolution[1326:903] Display 0: 1280x1024x32
<hedule wakeorpoweron "`date -v+1M "+%m/%d/%y %H:%M:%S"`" && shutdown -h nowConnection to talos-r4-snow-013.build.mozilla.org closed by remote host.
Connection to talos-r4-snow-013.build.mozilla.org closed.
Last login: Tue Dec 20 07:54:34 2011 from bm-vpn01.build.sjc1.mozilla.com
talos-r4-snow-013:~ cltbld$ screenresolution get
2011-12-20 07:54:56.168 screenresolution[123:903] starting screenresolution argv=screenresolution get
2011-12-20 07:54:56.194 screenresolution[123:903] Display 0: 1280x1024x32
talos-r4-snow-013:~ cltbld$ screenresolution list
2011-12-20 07:54:59.824 screenresolution[146:903] starting screenresolution argv=screenresolution list
Available Modes on Display 0
1x1x8 1x1x16 1x1x32 1x1x64
1x1x96 800x600x8 800x600x16 800x600x32
800x600x64 800x600x96 1024x768x8 1024x768x16
1024x768x32 1024x768x64 1024x768x96 1280x1024x8
1280x1024x16 1280x1024x32 1280x1024x64 1280x1024x96
1680x1050x8 1680x1050x16 1680x1050x32 1680x1050x64
1680x1050x96 1280x1024x32 talos-r4-snow-013:~ cltbld$
Sadly, it looks like this doesn't solve the problem.
| Assignee | ||
Comment 83•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #82)
> Sadly, it looks like this doesn't solve the problem.
but.... using a 3m time worked! I think its a capacitor or inductor holding charge between power cycles.
| Assignee | ||
Comment 84•14 years ago
|
||
The corresponding script has the contents:
~/mozilla/new-r4-reboot $ cat pmset-reboot.sh
#!/bin/bash
# Schedule a boot at 3, 10 and 60 minutes in case machine takes too long to
# shutdown.
# Clear out any jobs left over from previous boots. Doing this here as a
# safety net. We also don't know if there are adverse effects to having
# a wakeorpoweron event firing on a system in production
rm /Library/Preferences/SystemConfiguration/com.apple.AutoWake.plist
for i in 3 10 60 ; do
pmset schedule wakeorpoweron "$(date -v+${i}M '+%m/%d/%y %H:%M:%S')"
done
This will also need a change to /etc/sudoers to allow passwordless root on pmset-reboot.sh and to disallow passwordless '/sbin/reboot'. We'll also need to update tools to call pmset-reboot.sh instead of reboot.
Attachment #583255 -
Flags: review?(coop)
| Assignee | ||
Comment 85•14 years ago
|
||
(In reply to John Ford [:jhford] from comment #84)
> Created attachment 583255 [details] [diff] [review]
> changes to puppet to deploy the new reboot script
This is what's needed to test out the pmset scheduled boot.
Updated•14 years ago
|
Comment 86•14 years ago
|
||
Here's the etherpad gathering everything I can find on the topic so far (excluding raw data, but linking heavily):
https://etherpad.mozilla.org/2011-12-21-dongles-mtng
Comment 87•14 years ago
|
||
John -- do you want to schedule another meeting sometime next week to see where we're at with this? I can do the scheduling if you're prefer, but I thought I'd share the fun :)
| Assignee | ||
Comment 88•14 years ago
|
||
I have not seen any intermittent or consistent failure from any of the 75Ω dongle or dvi doctor machines. talos-r4-snow-015 has died, but it is unpingable and nothing suggests that this is a dongle or resolution issue.
We were seeing intermittent problems roughly once per day and have a multiple machines fail in the consistent state each day. I think we can call the 75Ω dongles and dvi doctors both a solution and we can say that the problem was that we had the incorrect resistance. I would expect that in the two weeks we've had the 75Ω dongles and three that we've had dvi doctors, that we'd have seen at least one consistent or intermittent failure.
I propose we deploy 75Ω dongles to all talos-r4 machines because they seem to fix the issues and are easier to deploy than dvi doctors. If we want to have another meeting to discuss this, I can organize one for later this week or early next, otherwise I'd like to move forward with the 75Ω dongles.
Comment 89•14 years ago
|
||
That's great news! I'll file a bug to get started on that deployment. I want to do a bit more experimentation in the interim in bug 713746, so I'll leave this open. Bug 714954 covers the 75Ω deployment.
Updated•14 years ago
|
| Assignee | ||
Comment 90•14 years ago
|
||
As noted in bug 713746#c15, a bug in slavealloc has caused the conclusion regarding 75Ω dongles in comment 88 to be based on insufficient data. Because of this, I am closing bug 714954 as INVALID. We can file a new bug if we decide to roll out the 75Ω dongles.
I am going to re-enable the 75Ω dongle machines in slavealloc. Aside from snow-018, all the dvi doctor machines were enabled in slavealloc, even though their comments were overwritten with incorrect data.
Comment 91•14 years ago
|
||
Remember I'm using snow-014. It won't start buildbot in its current state anyway, but it should still be disabled and commented as such.
Comment 92•14 years ago
|
||
Bug 713746 is mostly at an end now, I think.
Evidence suggests that a script similar to that given at the end of the bug will do the trick in production -- if it can't set the resolution, it cold-boots the box. As John said in comment 46, what we need is a way to work around the problem, since we've not been able find any way to actually solve it.
John, what do you think - should we put something like this into practice and see what it misses?
Updated•14 years ago
|
Attachment #583255 -
Flags: review?(coop) → review+
Comment 93•14 years ago
|
||
(In reply to Dustin J. Mitchell [:dustin] from comment #92)
> John, what do you think - should we put something like this into practice
> and see what it misses?
There are a bunch of different patches in flight that address parts of this issue:
* Attachment #583255 [details] [diff] (in this bug)
* Attachment #583849 [details] [diff] (bug 712750)
* the script in https://bugzilla.mozilla.org/show_bug.cgi?id=713746#c18
Do we need all these moving parts, or do some patches supercede others? Can we amalgamate some/all of these bugs, please?
| Assignee | ||
Comment 94•14 years ago
|
||
(In reply to Chris Cooper [:coop] from comment #93)
> (In reply to Dustin J. Mitchell [:dustin] from comment #92)
> > John, what do you think - should we put something like this into practice
> > and see what it misses?
>
> There are a bunch of different patches in flight that address parts of this
> issue:
>
> * Attachment #583255 [details] [diff] (in this bug)
This is a script that lets us do the cold reboot. Its purpose is to prevent pool size collapse, but does nothing to prevent failure.
> * Attachment #583849 [details] [diff] (bug 712750)
Extra reporting to help figure out what is going on.
> * the script in https://bugzilla.mozilla.org/show_bug.cgi?id=713746#c18
I assume this is a replacement for pmset-reboot.sh mentioned in patch 583255 (puppet changes attached to this bug)
> Do we need all these moving parts, or do some patches supercede others? Can
> we amalgamate some/all of these bugs, please?
The two r+'d patches are what we need if we continue trying to work around instead of solve.
| Assignee | ||
Comment 95•14 years ago
|
||
I would like to install DVI Doctors on all rev4 slaves. These devices have been running in production without issue since Dec 15, nearly a month. Not a single detected intermittent failure and none of these machines have been in the consistent failure state since the dvi doctors were installed.
I think spending the time and money on a known solution is a lot better than continuing to chase down the problem when after two months, we still don't know what the root problem is or even what triggers it.
Comment 96•14 years ago
|
||
Before we do that, let's run them through the testing regime in bug 713746.
| Assignee | ||
Comment 97•14 years ago
|
||
(In reply to Dustin J. Mitchell [:dustin] from comment #96)
> Before we do that, let's run them through the testing regime in bug 713746.
I have disabled the following slaves in slavealloc for your tests.
talos-r4-snow-011
talos-r4-snow-018
talos-r4-snow-023
When do you expect results from this test?
Comment 98•14 years ago
|
||
I think we've wrapped up startup resolution failures, which is what bug 713746 is about. I'd be happy with simply implementing this cold-boot-on-failure technique. I think that the cost is small enough that installing 75Ω dongles on all rev4's is acceptable too, without further research -- although as I stated above I'm not convinced it's necessary. I don't think we have any good reason to deploy DVI Doctors.
Comment 99•14 years ago
|
||
(In reply to Dustin J. Mitchell [:dustin] from comment #98)
> I think we've wrapped up startup resolution failures, which is what bug
> 713746 is about. I'd be happy with simply implementing this
> cold-boot-on-failure technique. I think that the cost is small enough that
> installing 75Ω dongles on all rev4's is acceptable too, without further
> research -- although as I stated above I'm not convinced it's necessary. I
> don't think we have any good reason to deploy DVI Doctors.
So I keep going back-and-forth on this depending on whether I've talked to server ops or jhford last.
DVI Doctors are known to work but are not cheap, and may entail further work reshuffling minis due to lack of space within racks. If we *need* to do this, I have no problem pulling the trigger on the purchase.
However we have software fixes in-hand that (according to https://bugzilla.mozilla.org/show_bug.cgi?id=713746#c27) will work around the known failure cases. Let's go ahead and deploy the software fixes before we incur the hardware costs.
| Assignee | ||
Comment 100•14 years ago
|
||
(In reply to Chris Cooper [:coop] from comment #99)
> (In reply to Dustin J. Mitchell [:dustin] from comment #98)
> > I think we've wrapped up startup resolution failures, which is what bug
> > 713746 is about. I'd be happy with simply implementing this
> > cold-boot-on-failure technique. I think that the cost is small enough that
> > installing 75Ω dongles on all rev4's is acceptable too, without further
> > research -- although as I stated above I'm not convinced it's necessary. I
> > don't think we have any good reason to deploy DVI Doctors.
>
> So I keep going back-and-forth on this depending on whether I've talked to
> server ops or jhford last.
>
> DVI Doctors are known to work but are not cheap, and may entail further work
> reshuffling minis due to lack of space within racks. If we *need* to do
> this, I have no problem pulling the trigger on the purchase.
>
> However we have software fixes in-hand that (according to
> https://bugzilla.mozilla.org/show_bug.cgi?id=713746#c27) will work around
> the known failure cases. Let's go ahead and deploy the software fixes before
> we incur the hardware costs.
This is not a fix for the dongle issues. This is a utility that cleans up after a failure so we don't lose the entire slave pool. The only thing that fixes this issue is the DVI Dongle. The utility from 713746 is not a fix, it does not prevent failure and still allows resolution issues to cause developer impacting test failures.
Comment 101•14 years ago
|
||
You'll recall, from that meeting, that we agreed to focus on the startup problems, both in hopes we could fix it, and that we could learn something about the behavior of the minis. We can argue the semantics of "fix", but the point is that the script described in bug 713746 makes the startup problem go away, with any sort of dongle. We've also learned some things about the resolution behavior -- in particular, that it can change without warning.
Nothing in those tests indicates that the change can only occur on (warm) reboot, so it's certainly reasonable to think that the failures occur any time. What data do we have to support that? How significant is the evidence that the DVI Doctors fix runtime failures?
I only know of the three instances of runtime failures from November and early December cited in
https://etherpad.mozilla.org/2011-12-21-dongles-mtng
have there been more? I'd like to know how many tests have been run on each kind of dongle over a given time period, and how many runtime failures have occurred for each kind of dongle.
We planned, in the meeting, to begin gathering more comprehensive data. How is that going?
| Assignee | ||
Comment 102•14 years ago
|
||
During a meeting today we decided to go ahead with DVI Doctors. I have filed bug 720006 to request that we install dvi doctors on all the machines.
Marking this bug as resolved because we now know what the solution to these problems are. Please track future work in bug 720006.
Status: NEW → RESOLVED
Closed: 14 years ago
Resolution: --- → FIXED
Comment 103•14 years ago
|
||
My apologies for being out of the loop here. Where can I see the data that was collected here:
(In reply to Dustin J. Mitchell [:dustin] from comment #101)
> We planned, in the meeting, to begin gathering more comprehensive data. How
> is that going?
That lead to this?:
(In reply to John Ford [:jhford] from comment #102)
> During a meeting today we decided to go ahead with DVI Doctors. I have
> filed bug 720006 to request that we install dvi doctors on all the machines.
Comment 104•14 years ago
|
||
There was no additional data, which is a shame.
However, the combination of
* no runtime failures on DVI doctors in production
* several (John didn't have a number) failures on both 75Ω and 100Ω dongles
* DVI Doctors' good behavior in my tests of the startup failures, and
* no discernable difference in startup failure behavior between 75Ω and 100Ω
suggests (say, p<0.15) that DVI Doctors *do* fix the problems (not surprisingly), whereas (more surprisingly) dongles do not. The software fix I created for startup failures cannot address runtime failures.
While I'm unhappy with the lack of meaningful data on runtime failures here, I don't think that additional (experimental) data will appreciably alter the chosen remedy.
Updated•12 years ago
|
Component: Server Operations: RelEng → RelOps
Product: mozilla.org → Infrastructure & Operations
You need to log in
before you can comment on or make changes to this bug.
Description
•