Closed
Bug 1318890
Opened 9 years ago
Closed 9 years ago
Archiver broken by relengapi failure after TCW
Categories
(Release Engineering :: General, defect)
Release Engineering
General
Tracking
(Not tracked)
RESOLVED
FIXED
People
(Reporter: aryx, Unassigned)
References
Details
Attachments
(8 files)
Buildbot jobs fail.
Closed trees for this, affects also Try.
https://treeherder.mozilla.org/#/jobs?repo=mozilla-inbound&revision=9f8700acec463beb0f1d37bcd3946a69a152f7a3
E.g. https://treeherder.mozilla.org/logviewer.html#?job_id=39507020&repo=mozilla-inbound
2016-11-19 13:55:43,467 truncating revision to first 12 chars
2016-11-19 13:55:43,467 Setting DEBUG logging.
2016-11-19 13:55:43,467 attempt 1/10
2016-11-19 13:55:43,467 Getting archive location from https://api.pub.build.mozilla.org/archiver/hgmo/integration/mozilla-inbound/9f8700acec46?&preferred_region=us-west-2&suffix=tar.gz&subdir=testing/mozharness
2016-11-19 13:55:44,176 attempt 1/10
2016-11-19 13:55:44,749 current task status: no status available at this point. state: PENDING
2016-11-19 13:55:44,750 sleeping for 10.00s (attempt 1/10)
2016-11-19 13:55:54,750 attempt 2/10
2016-11-19 13:55:55,296 current task status: no status available at this point. state: PENDING
2016-11-19 13:55:55,296 sleeping for 16.00s (attempt 2/10)
2016-11-19 13:56:11,296 attempt 3/10
2016-11-19 13:56:11,841 current task status: no status available at this point. state: PENDING
2016-11-19 13:56:11,841 sleeping for 25.00s (attempt 3/10)
2016-11-19 13:56:36,841 attempt 4/10
2016-11-19 13:56:37,473 current task status: no status available at this point. state: PENDING
2016-11-19 13:56:37,473 sleeping for 36.50s (attempt 4/10)
2016-11-19 13:57:13,973 attempt 5/10
2016-11-19 13:57:14,513 current task status: no status available at this point. state: PENDING
2016-11-19 13:57:14,513 sleeping for 55.75s (attempt 5/10)
2016-11-19 13:58:10,263 attempt 6/10
2016-11-19 13:58:10,903 current task status: no status available at this point. state: PENDING
2016-11-19 13:58:10,903 sleeping for 84.62s (attempt 6/10)
2016-11-19 13:59:35,528 attempt 7/10
2016-11-19 13:59:36,138 current task status: no status available at this point. state: PENDING
2016-11-19 13:59:36,138 sleeping for 101.00s (attempt 7/10)
2016-11-19 14:01:17,138 attempt 8/10
2016-11-19 14:01:17,770 current task status: Task has expired from pending for too long. Re-creating task. state: RETRY
2016-11-19 14:01:17,770 sleeping for 100.00s (attempt 8/10)
2016-11-19 14:02:57,770 attempt 9/10
2016-11-19 14:02:58,392 current task status: no status available at this point. state: PENDING
2016-11-19 14:02:58,392 sleeping for 101.00s (attempt 9/10)
2016-11-19 14:04:39,392 attempt 10/10
2016-11-19 14:04:39,984 current task status: no status available at this point. state: PENDING
2016-11-19 14:04:39,984 Automation Error: task state did not equal SUCCESS
2016-11-19 14:04:39,984 Check archiver logs for errors. Task status: no status available at this point. Task state: PENDING
program finished with exit code 3
Updated•9 years ago
|
Component: General Automation → Buildduty
QA Contact: catlee → bugspam.Callek
log from running ./upsdate on relengwebadm as recommended in #releng -- unfortunately, not all dependencies are pinned and some updates failed.
[14:58] <@ catlee>| I think that's relengapi
[14:59] <@ catlee>| Maybe needs a kick?
[16:07] < sal>| catlee: still closed?
[16:37] <@nthomas|awa>| I agree with c.atlee. Not at a computer user right now though
[16:42] <@nthomas|awa>| Docs at https://m.wiki.mozilla.org/ReleaseEngineering/Applications/RelengAPI
[16:42] <@nthomas|awa>| relengadm
[16:51] * | hwine-ooo jumps on that now that Im near a computer
[17:01] < hwine> | oh wonderful - the doc'd proceedure took down relengapi :/
Comment 2•9 years ago
|
||
relengapi has had service restored. The automation that builds the virtualenv failed badly and needs to be fixed still. We've stabilized the cluster for the weekend and will attend to it on Monday.
We have jobs (completed prior to TCW) being able to be re-run successfully, so declaring current machine state okay.
We (releng) are cleaning up some of the broken jobs before reopening trees.
| Reporter | ||
Updated•9 years ago
|
Severity: blocker → normal
As part of just recording a "known working state", we'd like to get a "pip freeze" and/or "rpm -qa" attached for each of the celery servers (one problem area).
:ericz - can you generate that for whatever you had to fix on Saturday, please?
Flags: needinfo?(eziegenhorn)
Comment 6•9 years ago
|
||
The virtualenv is shared across them all supposedly, though it seemed there were problems with syncing when I was putting out the fire over the weekend. pip freeze isn't working at the moment but here is the requirements that *should've* been installed in the virtualenv:
(virtualenv)[eziegenhorn@relengwebadm.private.scl3 relengapi]$ cat requirements.txt
relengapi[ldap]==3.3.8
Sphinx==1.3.1
MySQL-python==1.2.5
ConcurrentLogHandler==0.9.1
# For Monitoring
newrelic==2.46.0.37
# For Issue 347
nose
# for monitoring (latest is OK)
flower
On the working virtualenv that I monkey patched into the web and celery nodes on Saturday are these python 2.7 packages:
alabaster-0.7.3
alembic-0.7.7
anyjson-0.3.3
argparse-1.2.1
Babel-1.3
backports.ssl_match_hostname-3.4.0.2
billiard-3.3.0.23
blinker-1.3
boto-2.38.0
bzrest-1.1
celery-3.1.22
ConcurrentLogHandler-0.9.1
croniter-0.3.5
docutils-0.11
elasticache_auto_discovery-0.0.5
Flask-0.10.1
Flask_BrowserID-0.0.4
Flask_Login-0.3.0
flower-0.8.3
furl-0.4.5
ipaddr-2.1.11
IPy-0.75
itsdangerous-0.24
Jinja2-2.7.1
kombu-3.0.34
logging_tree-1.6
Mako-1.0.1
MarkupSafe-0.18
mohawk-0.3.2.1
mozdef_client-1.0.4
MySQL_python-1.2.5
nose-1.3.7
orderedmultidict-0.7.5
python_dateutil-1.5
python_ldap-2.4.15
python_memcached-1.53
pytz-2015.6
relengapi-3.3.8
requests_futures-0.9.5
setuptools-4.0.1
simplegeneric-0.8
simplejson-3.3.0
slugid-1.0.7
snowballstemmer-1.2.0
Sphinx-1.3.1
SQLAlchemy-1.0.3
taskcluster-0.2.0
tornado-4.2.1
WebOb-1.2.3
Werkzeug-0.9.3
wrapt-1.8.0
WSME-0.6.1
More info coming soon.
Flags: needinfo?(eziegenhorn)
Full requirements from properly functioning web1. Identical with file from web2.
Full requirements from properly functioning web2. Identical with file from web1.
requirements from virtualenv in relengwebadm:/data/releng/www/relengapi/virtualenv
Does not match what is in production.
Comment 10•9 years ago
|
||
requirements from virtualenv in relengwebadm:/data/releng/src/relengapi/virtualenv
Does not match what is in production, or what is in 'www' directory
Comment 11•9 years ago
|
||
The update script which was used for code updates and general service restarts does no error checking and if pip fails as it did over the weekend, it'll deploy a broken virtualenv to the webheads and celery nodes and restart services, causing the site to be somewhere from not right to completely not working.
Additionally the webheads and celery nodes pull all of the code every 5 minutes from relengwebadm.
I still don't know what went wrong with the virtualenv on Saturday.
Also, I've noticed that celery2.srv.releng.scl3 is not included in the deploy scripts on relengwebadm. It is pulling code in the same way every 5 minutes, but a deploy run on relengwebadm does nothing to it.
I recommend:
1. Build error checking into the update script so that it refuses to deploy when errors are encountered
2. Turn off the every 5 minutes auto-updating unless absolutely needed
3. Adding celery2 to the commander configs on relengwebadm so it receives code pushes
4. Fixing the virtualenv
5. Verify the code deployment infrastructure works once the virtualenv is fixed. I had a lot of trouble with this on Saturday but likely that was just due to the pressure of fire fighting as I don't think there have been any changes there.
Most of this could be tested on stage.
Comment 12•9 years ago
|
||
I've improved the update script on relengapi stage so that it rebuilds the virtualenv from scratch every run and bails out with a helpful error message if it encounters any problems. I had to add unlisted dependencies to requirements.txt to get everything to work correctly as well. Hal, when can I roll out this improved update script and virtualenv and commander config fixes as per comment 11 to prod?
Flags: needinfo?(hwine)
Updated•9 years ago
|
Component: Buildduty → General Automation
QA Contact: bugspam.Callek → catlee
Summary: Buildbot jobs fail → Archiver broken by relengapi failure after TCW
Comment 14•9 years ago
|
||
sorted list from relengwebadm virtualenv
Comment 15•9 years ago
|
||
(In reply to Eric Ziegenhorn :ericz from comment #12)
> Hal, when can
> I roll out this improved update script and virtualenv and commander config
> fixes as per comment 11 to prod?
Assuming the testing mentioned in comment 11 passed, I see one issue to resolve before we roll out. The roll out would be a good candidate for a Wed Dec 21 morning deploy.
The remaining issue is how to handle the requirements.txt file. Supposedly, that file is:
a) kept in source control[1]
b) manually updated on relengwebadm box[2]
However, I can not reconcile the current version (3.3.8)[1][3] with what is currently specified (and working) in the virtualenv on the box.[4]
Based on the expected lifetime of the service, I suspect the correct answer is to:
i) keep using the current requirements.txt on the production box
ii) only deploy new code based on the 3.3.8 tag (substantial changes have been landed afterwards, but not deployed as far as I can see)
setting ni on :dustin & :callek for comments on this plan
[1] https://github.com/mozilla/build-relengapi/blob/relengapi-3.3.8/requirements.txt (62 lines)
[2] instructions in 6th paragraph of section 2 of https://wiki.mozilla.org/ReleaseEngineering/How_To/Update_RelengAPI
[3] sorted version of [1] for better diff as attachment 8818997 [details]
[4] sorted version of output of '/data/releng/src/relengapi/virtualenv/bin/pip freeze' (64 lines) as attachment 8818999 [details]
Flags: needinfo?(hwine)
Flags: needinfo?(dustin)
Flags: needinfo?(bugspam.Callek)
Comment 16•9 years ago
|
||
(In reply to Hal Wine [:hwine] (use NI) from comment #15)
> (In reply to Eric Ziegenhorn :ericz from comment #12)
> > Hal, when can
> > I roll out this improved update script and virtualenv and commander config
> > fixes as per comment 11 to prod?
>
> Assuming the testing mentioned in comment 11 passed, I see one issue to
> resolve before we roll out. The roll out would be a good candidate for a Wed
> Dec 21 morning deploy.
>
> The remaining issue is how to handle the requirements.txt file. Supposedly,
> that file is:
> a) kept in source control[1]
> b) manually updated on relengwebadm box[2]
>
> However, I can not reconcile the current version (3.3.8)[1][3] with what is
> currently specified (and working) in the virtualenv on the box.[4]
>
> Based on the expected lifetime of the service, I suspect the correct answer
> is to:
> i) keep using the current requirements.txt on the production box
> ii) only deploy new code based on the 3.3.8 tag (substantial changes have
> been landed afterwards, but not deployed as far as I can see)
>
> setting ni on :dustin & :callek for comments on this plan
>
> [1]
> https://github.com/mozilla/build-relengapi/blob/relengapi-3.3.8/requirements.
> txt (62 lines)
Looks to only have landed with the commit message "Project Gardening" by rgarbas
> [2] instructions in 6th paragraph of section 2 of
> https://wiki.mozilla.org/ReleaseEngineering/How_To/Update_RelengAPI
> [3] sorted version of [1] for better diff as attachment 8818997 [details]
> [4] sorted version of output of
> '/data/releng/src/relengapi/virtualenv/bin/pip freeze' (64 lines) as
> attachment 8818999 [details]
The latter output is likely taken from the on-relengweb host, with specifically:
https://github.com/mozilla/build-relengapi/blob/relengapi-3.3.8/setup.py
What the basic design used to be is this:
* a *not* version controlled requirements.txt
-- In here was iirc `relengapi['ldap']==<version>` and a few other needed things not in version controll there (e.g. new relic)
* This used requirements.txt was actually used to install the newer dep of relengapi, and then based on its setup.py would pull in newer (or not-newer) deps as necessary.
* ./update.sh would run pip against that requirements.txt, and update the venv, then sync the venv update out to the webheads (and the celery head) and restart services as needed.
---
I had previously investigated upgrading the list of packages on the host, with the likes of Bug 1084012, prior to rgarbas taking over.
I think the overall plan is sound, but we should probably reconcile the plan with how we used to deploy, ideally keeping the older copies of dep packages to avoid potential extra issues
Flags: needinfo?(bugspam.Callek)
Comment 17•9 years ago
|
||
Yes sorting out what should and should not be in requirements.txt is the trickiest remaining part. What I did in stage was a somewhat laborious working out what it needs by trying to run the app, see what's broken/missing, add that missing library and then try running it again. It's not ideal but it isn't uncommon either and should work in lieu of someone knowing the actual requirements. Hopefully someone knows what the canonical requirements list is but I can work through it in this manner as I did with stage if needed.
Comment 18•9 years ago
|
||
For comparison, the requirements.txt that seems to be working in stage is much shorter.
Comment 19•9 years ago
|
||
Comment on attachment 8819009 [details]
relengapi-stage requirements.txt
I do find it odd that we needed to manually install stuff for wsme.
Only thing I can think of is that we didn't re-parse what was in relengapi[ldap]==3.3.8 in what could have caused that...
Comment 20•9 years ago
|
||
(In reply to Justin Wood (:Callek) from comment #16)
> (In reply to Hal Wine [:hwine] (use NI) from comment #15)
> > [1]
> > https://github.com/mozilla/build-relengapi/blob/relengapi-3.3.8/requirements.
> > txt (62 lines)
>
> Looks to only have landed with the commit message "Project Gardening" by
> rgarbas
Regardless of commit message, it is what was "shipped" in the 3.3.8 release tarball. Whether or not that requirements file is referenced by any installation code is a different issue.
>
> > [4] sorted version of output of
> > '/data/releng/src/relengapi/virtualenv/bin/pip freeze' (64 lines) as
> > attachment 8818999 [details]
>
> The latter output is likely taken from the on-relengweb host, with
> specifically:
>
> https://github.com/mozilla/build-relengapi/blob/relengapi-3.3.8/setup.py
No, the latter output is the result of the command I gave, and would only indirectly have a relationship with the setup.py contents.
>
> What the basic design used to be is this:
>
> * a *not* version controlled requirements.txt
> -- In here was iirc `relengapi['ldap']==<version>` and a few other needed
> things not in version controll there (e.g. new relic)
> * This used requirements.txt was actually used to install the newer dep of
> relengapi, and then based on its setup.py would pull in newer (or not-newer)
> deps as necessary.
> * ./update.sh would run pip against that requirements.txt, and update the
> venv, then sync the venv update out to the webheads (and the celery head)
> and restart services as needed.
So, I think the TCW experience showed that this approach is no longer deterministic enough, given our limited in depth knowledge of the system.
> I think the overall plan is sound, but we should probably reconcile the plan
> with how we used to deploy, ideally keeping the older copies of dep packages
> to avoid potential extra issues
Per IRC convo - I think we have agreement that we would only re-deploy to address a severe bug or security issue in the codebase. Any 3rd party sec/bug would be handled by onhost installation of packages into the venv. There are no know business reasons to upgrade this code base, as it is now legacy.
So, :ericz, you have the goahead to rollout the changes in comment 11. Please coordinate with buildduty per standard practice.
Flags: needinfo?(hwine)
Flags: needinfo?(eziegenhorn)
Flags: needinfo?(dustin)
Comment 21•9 years ago
|
||
Update: I attempted fixing this yesterday and made some progress but in pulling web1.releng.webapp.scl3 out of the load balancer we started seeing alerts and it appears pypi.pub not being able to handle the load running on just 1 of the 2 web heads. I dug deeply into why this was the case and came up empty handed, the single webhead seemed lightly loaded in terms of network, cpu, disk i/o and memory. There was a suggestion that mod_wsgi process tuning might help by pypi does not use mod_wsgi and the prefork MPM configuration looks sane and appropriate to me. Hal decided to punt this to after the holidays given the time constraints and difficulty. At that point we have 3 options that I see:
1) Proceed without taking a webhead out of the load balancer. This is possible but slightly riskier.
2) Add another web head to the releng web cluster.
3) Dive deeper into Apache/Zeus/etc for tuning to try and raise utilization of the existing webheads.
At this point I'm leaning toward 2) adding another web head as it's moderately simpler and should help with capacity problems that we occasionally see on this cluster as well as this maintenance we're trying to accomplish.
Flags: needinfo?(eziegenhorn)
Comment 22•9 years ago
|
||
Agree regarding (2) as webheads often come out for other sorts of maintenance as well.
Comment 23•9 years ago
|
||
(In reply to Eric Ziegenhorn :ericz from comment #21)
> Update: I attempted fixing this yesterday and made some progress but in
> pulling web1.releng.webapp.scl3 out of the load balancer we started seeing
> alerts and it appears pypi.pub not being able to handle the load running on
> just 1 of the 2 web heads. I dug deeply into why this was the case and came
> up empty handed, the single webhead seemed lightly loaded in terms of
> network, cpu, disk i/o and memory. There was a suggestion that mod_wsgi
> process tuning might help by pypi does not use mod_wsgi and the prefork MPM
> configuration looks sane and appropriate to me. Hal decided to punt this to
> after the holidays given the time constraints and difficulty. At that point
> we have 3 options that I see:
>
> 1) Proceed without taking a webhead out of the load balancer. This is
> possible but slightly riskier.
> 2) Add another web head to the releng web cluster.
> 3) Dive deeper into Apache/Zeus/etc for tuning to try and raise utilization
> of the existing webheads.
>
> At this point I'm leaning toward 2) adding another web head as it's
> moderately simpler and should help with capacity problems that we
> occasionally see on this cluster as well as this maintenance we're trying to
> accomplish.
In addition to option 2, we can also *probably* increase caching for pypi mirror itself.
Comment 24•9 years ago
|
||
Question: is the new webhead in the pool now? (I can't remember)
Note on "fixing" process for other 2 webheads:
- after removing from zlb pool, apache must be graceful-stop
- coordination with #releng to get ack it's an okay time
Comment 25•9 years ago
|
||
(In reply to Hal Wine [:hwine] (use NI) from comment #24)
> - after removing from zlb pool, apache must be graceful-stop
Forgot to add "why" -- one of the web apps has a r/w connection to the buildbot databases - any pending transactions need to complete cleanly.
Comment 26•9 years ago
|
||
The virtualenv was fixed on Friday morning and everything looks fine now. Deploys are working well, puppet is happy, the update script has been improved. Calling this done.
Status: NEW → RESOLVED
Closed: 9 years ago
Resolution: --- → FIXED
| Assignee | ||
Updated•8 years ago
|
Component: General Automation → General
You need to log in
before you can comment on or make changes to this bug.
Description
•