Closed Bug 1369686 Opened 9 years ago Closed 8 years ago

OS X taskcluster workers for NSS

Categories

(Taskcluster :: Workers, enhancement)

enhancement
Not set
normal

Tracking

(Not tracked)

RESOLVED FIXED

People

(Reporter: pmoore, Unassigned)

References

Details

Previously OS X builds ran on buildbot infrastructure, and for a while NSS OS X builds have been cross-compiled on linux, however this is causing some problems. The NSS devs would like to run the OS X builds natively on OS X, which would mean using workers in scl3 as we have things at the moment. I think it would make most sense therefore for them to use releng-managed machines. Initially it will be just builds, but ideally later also tests. There are few pushes per day on nss-try and nss, but when there is a push, having a pool of workers available would be useful (e.g. 20-30 worker capacity). A dedicated pool probably doesn't make much sense if resources are constrained, and therefore it would probably make sense in terms of efficiency to go on a shared pool. I believe at the moment we only have a single worker type per platform. Therefore, the first decision, if releng agrees to host the workers, is whether NSS should use the existing gecko worker type, or have their own worker type. If using gecko, are the security implications? If their own, deciding what pool size will be key - too big, wasted idle resources. Too small, and there will be long end-to-end times getting test results. Maybe shared pool for try, and a dedicated constrained worker pool for non-try? Catlee, what are your thoughts?
Flags: needinfo?(catlee)
Summary: OS X workers for NSS → OS X taskcluster workers for NSS
Depends on: 1358533
Chris, can you give a timeline for this? We'd need them soonish as we had to disabled some of the old buildbots.
This is not a good long-term plan. We are still hoping to use cross-compiled builds for OSX for Firefox, and so are not planning to have any OSX build capacity by the end of the year. The machines in scl3 are planned to be decommissioned. What are the issues around cross-compiling NSS? Also, AIUI, we currently have no production-ready OSX taskcluster workers in SCL3. We've done some testing there, but aren't yet committed to migrating OSX builds to those workers.
Flags: needinfo?(catlee)
(In reply to Chris AtLee [:catlee] from comment #2) > This is not a good long-term plan. Can you elaborate on why that is? > What are the issues around cross-compiling NSS? GYP is giving us a hard time and seems not have been written with that feature in mind. We'd love to not have to patch/extend it. > Also, AIUI, we currently have no production-ready OSX taskcluster workers in > SCL3. We've done some testing there, but aren't yet committed to migrating > OSX builds to those workers. We'd be more than happy to be guinea pigs and help test OS X workers.
(In reply to Chris AtLee [:catlee] from comment #2) > Also, AIUI, we currently have no production-ready OSX taskcluster workers in > SCL3. We've done some testing there, but aren't yet committed to migrating > OSX builds to those workers. We've now migrated all OS X tests from buildbot to taskcluster (generic-worker) in scl3. So this part, at least is solved. Catlee, does comment 3 change the landscape in any way for you? Other options I can think of (if releng/relops hosting is not possible) are: 1) using something like https://www.macstadium.com/ 2) getting a couple of mac minis for the berlin office to at least stand up builds, even if this doesn't solve the "tests" problem yet
Flags: needinfo?(catlee)
NSS and firefox are separate projects and will not use the same infrastructure. I've emailed Tim to talk about requirements and solutions.
Status: NEW → RESOLVED
Closed: 9 years ago
Flags: needinfo?(catlee)
Resolution: --- → INVALID
The bug title is "OS X taskcluster workers for NSS" and doesn't mention that this has to be on the same infrastructure as firefox. We should keep this topic in the public domain since we're an open source organisation, and there might be open source NSS devs/testers watching this bug, wondering what the CI strategy will be on macOS. Thanks!
Status: RESOLVED → REOPENED
Resolution: INVALID → ---
This bug will be moved under the Taskcluster component and we will track the strategy/progress for supporting a worker for taskcluster for NSS. There is no CI strategy as of today, but once there is infrastructure to support it, it can be documented in this bug along with a strategy going forward.
Component: Platform Support → Worker
Product: Release Engineering → Taskcluster
QA Contact: catlee
Talked to Amy last week and I'll work with her to get us a few MacStadium machines. Then I'll work with Pete in this bug to get the docker-worker running and everything set up.
Just so this is on the collective radar, we now have a hard deadline to get all of our gear out of our SCL3 colo by September of next year.
NSS team have their own workers in macstatdium.com now and are running jobs on nss and nss-try in tier 3.
Status: REOPENED → RESOLVED
Closed: 9 years ago8 years ago
Resolution: --- → FIXED
OK, to be completely clear: does this mean that we can shut down and decommission NSS hardware running in SCL3?
Flags: needinfo?(ttaubert)
(In reply to Mike Hoye [:mhoye] from comment #11) > OK, to be completely clear: does this mean that we can shut down and > decommission NSS hardware running in SCL3? I think we're ready to shut down and decommission. Kai, is there anything (data?) left on these machines you want to copy before we do that? Or can we just go ahead?
Flags: needinfo?(ttaubert) → needinfo?(kaie)
We have weekly cronjob running on nss-vm-centos6-1 that performs fuzzing tests with the NISCC test data. It was a bit of work to get that set up on that machine. This job is currently running only on that machine, and on another machine hosted inside the private Red Hat network. The environment isn't fully documented or backed up, so these two machines served each as the backup of the other, should it ever become necessary to setup an additional machine. In an ideal world it would be nice to migrate the /home/niscc.sqsh compressed image and the niscc user account from that machine, to another Linux VM, and keep it running somewhere. It's up to Mozilla if you want to do that or not. If not, you'll rely on me to not loose the environment we currently have at Red Hat, and potentially set it up on a new additional system.
Flags: needinfo?(kaie)
Besides what I mentioned in comment 13, I don't need to copy data, the other machines assigned to NSS in scl3 are no longer required by me.
Per an email conversation, Kai, you asked if we could ship these machines to the Berlin office. I'm happy to arrange that. If you're OK with them being powered down and shipped, we'll do that and you can document and back them up at your convenience. All I care about is getting them out of their colo. Maybe rsync them to somewhere beforehand just so you've got more than no backups during shipping.
(In reply to Kai Engert (:kaie:) from comment #13) > It's up to Mozilla if you want to do that or not. If not, you'll rely on me > to not loose the environment we currently have at Red Hat, and potentially > set it up on a new additional system. I'm okay with RedHat maintaining the only system with those tests. We can investigate later if and how we would want to revive that. (In reply to Mike Hoye [:mhoye] from comment #15) > Per an email conversation, Kai, you asked if we could ship these machines to > the Berlin office. > > I'm happy to arrange that. If you're OK with them being powered down and > shipped, we'll do that and you can document and back them up at your > convenience. All I care about is getting them out of their colo. That honestly doesn't sound like a good use of our time and money. We should rather invest into easily scalable testing with e.g. Taskcluster and not keep old machines around that we have to maintain ourselves. Franziskus has been working for a while on a fuzzer that generates certificates and we can integrate with TC and OSS-Fuzz. This is way more useful than a huge static set of certificates and a test runner that is quite difficult to set up.
OK, I'm going to have that hardware decommed shortly and shelved somewhere reasonably secure for six months. We can put them back in service if we need to during that time, and if not I'll have them scrubbed and disposed of. Thanks, everyone.
Component: Worker → Workers
You need to log in before you can comment on or make changes to this bug.