Closed Bug 612792 Opened 15 years ago Closed 15 years ago

devicemanager dies on "address already in use"

Categories

(Testing :: General, defect)

ARM
Android
defect
Not set
normal

Tracking

(Not tracked)

RESOLVED FIXED

People

(Reporter: mozilla, Assigned: bear)

References

Details

Attachments

(4 files, 4 obsolete files)

This might not be a fair use case, as I did an adb reboot in the middle of updateApp() to avoid wasting more time testing than I needed to during a test run that I knew would fail. The very next run died with files are validated updateApp using command: updt org.mozilla.fennec /mnt/sdcard/tests/fennec-4.0b2pre.en-US.eabi-arm.apk 10.250.48.9 30000 Creating server with 10.250.48.9:30000 Traceback (most recent call last): File "../../sut_tools/installApp.py", line 41, in <module> if dm.updateApp(target, processName='org.mozilla.fennec', ipAddr="10.250.48.9"): File "/builds/sut_tools/devicemanager.py", line 768, in updateApp callbacksvr = callbackServer(ip, port) File "/builds/sut_tools/devicemanager.py", line 812, in __init__ self.server = SocketServer.TCPServer((ip, port), self.myhandler) File "/tools/python-2.6.5/lib/python2.6/SocketServer.py", line 400, in __init__ self.server_bind() File "/tools/python-2.6.5/lib/python2.6/SocketServer.py", line 411, in server_bind self.socket.bind(self.server_address) File "<string>", line 1, in bind socket.error: [Errno 98] Address already in use I'm not sure whether we need to add more smarts on our end or on the devicemanager end.
The run after this seems to be better, though... ?
Bear: this may become an issue with 14-44 devices all on bm-foopy trying to open port 30000. You may need to have a starting port # associated with each tegra (e.g. tegra-001 starts at 30010, tegra-002 starts at 30020, etc.) in the bm-foopy slave side.
Blocks: 561908
Kept running into this. This workaround should reduce the frequency until we fix it for real: # This is retarded but until bug 612792 is fixed ... port = random.randint(30000, 50000) # We need to not hardcode this IP. Breaks w/out it though if dm.updateApp(target, processName='org.mozilla.fennec', ipAddr="10.250.48.9", port=port):
Still an issue; I have - def updateApp(self, appBundlePath, processName=None, destPath=None, ipAddr=None, port=None): + def updateApp(self, appBundlePath, processName=None, destPath=None, ipAddr=None, port=30000): status = None cmd = 'updt ' if (processName == None): @@ -766,10 +862,7 @@ if (destPath): cmd += " " + destPath - if port: - ip, port = self.getCallbackIpAndPort(ipAddr, port) - else: - ip, port = self.getCallbackIpAndPort(ipAddr, 30000) + ip, port = self.getCallbackIpAndPort(ipAddr, port) in devicemanager to work around this (and set the port in installApp.py).
I think we have at least 3 layers of talos patches on foopy, alone; we need a patch queue or a user repo =P
If I understand the problem correctly (something similar was happening with the Python test agent), the solution is to do this SocketServer.TCPServer.allow_reuse_address = True at some point before creating the TCPServer in callbackServer.
Blocks: 610600
remove ip/port from updateApp() call since now dialback is handled by clientproxy.py change (bug 618363) Also refactored/cleaned-up reboot.py and cleanup.py
Assignee: nobody → bear
Attachment #501607 - Flags: review?(aki)
Two things, after skimming: * are we planning on eventually having a real port fix, rather than hoping that lightning doesn't strike twice in our randint? (real would be having a set of ports reserved per tegra, for instance) * are we planning on replacing the sharks error message with something more useful?
we don't need to have any port or IP now, randomly generated or whatever. if I missed one then I need to rework the patch to remove it. yea, I can put in something a bit more professional sounding :)
implemented suggestions, removed port/ip part in reboot.py I missed
Attachment #501607 - Attachment is obsolete: true
Attachment #501835 - Flags: review?(aki)
Attachment #501607 - Flags: review?(aki)
Comment on attachment 501835 [details] [diff] [review] patch to remove ip/port from dm calls I think these look ok. Do we exit out appropriately if we hit an error? Buildbot won't know to halt if we don't... we might want to add one to cleanup if not. Maybe that's handled by the errors->take device out of pool.
Attachment #501835 - Flags: review?(aki) → review+
This is the latest versions of these tools. the ip/port for dm calls had to be retained because removing it introduced an even gnarlier issue.
Attachment #501835 - Attachment is obsolete: true
Attachment #504695 - Flags: review?(aki)
Attachment #504695 - Flags: review?(aki) → review+
first try to get the proxy address from the environment and if not found determine what the IP address is of the server the script is run on
Attachment #504963 - Flags: review?(aki)
fixed typo in prev patch and also remembered to tell bugzilla it's a patch
Attachment #504963 - Attachment is obsolete: true
Attachment #504981 - Flags: review?(aki)
Attachment #504963 - Flags: review?(aki)
Comment on attachment 504981 [details] [diff] [review] determine proxy address instead of hardcoding it Already expressed my concerns with adding a mozilla.com:80 dependency to the scripts. We can change this to hardcode our little web cluster, or try to parse ifconfig, or something else. I think requiring CP_IP is the best route for now.
Attachment #504981 - Flags: review?(aki) → review+
adjust code to take advantage of new return values
Attachment #505713 - Flags: review?(aki)
Comment on attachment 505713 [details] [diff] [review] adjust helper scripts to new devicemanager.py return values Looks like we can hold off on updating cleanup.py til the removeDir() bug is addressed.
Attachment #505713 - Flags: review?(aki) → review+
(In reply to comment #17) > Comment on attachment 505713 [details] [diff] [review] > adjust helper scripts to new devicemanager.py return values > > Looks like we can hold off on updating cleanup.py til the removeDir() bug is > addressed. I made a local change to not check the removeDir() result so we can see how they do overnight with the other changed helper code
latest version adjusts to new return values and also includes standard error slug to allow the post step error parser to flag failures as infrastructure events
Attachment #505713 - Attachment is obsolete: true
Attachment #505973 - Flags: review?(aki)
Attachment #505973 - Flags: review?(aki) → review+
committed changeset 1074:7ed7ea2c708c
simple, ugly but works: set error.flg and then sleep() so that clientproxy's loop has a fighting chance of seeing it before buildmaster hands the slave 8 jobs in 2 minutes which all turn red/purple
Attachment #506952 - Flags: review?(aki)
Comment on attachment 506952 [details] [diff] [review] detect tegra/dm error and set semaphore to let clientproxy know Thinking maybe we'll want a timestamp in the .flg files at some point, but then again mtime may suffice.
Attachment #506952 - Flags: review?(aki) → review+
(In reply to comment #22) > Comment on attachment 506952 [details] [diff] [review] > detect tegra/dm error and set semaphore to let clientproxy know > > Thinking maybe we'll want a timestamp in the .flg files at some point, but then > again mtime may suffice. The routine to set the flag already has a parameter to insert text - just, at the time, couldn't think of anything useful to insert.
committed changeset 1087:5246670fa3b0
Aki - would like to close this as we are now paired up with how devicemanager is working.
Status: NEW → RESOLVED
Closed: 15 years ago
Resolution: --- → FIXED
Component: New Frameworks → General
You need to log in before you can comment on or make changes to this bug.

Attachment

General

Created:
Updated:
Size: