Closed Bug 1452837 Opened 8 years ago Closed 8 years ago

Online News v2 Data growing large for analysis cluster size

Categories

(Data Platform and Tools Graveyard :: Operations, enhancement)

enhancement
Not set
normal

Tracking

(Not tracked)

RESOLVED FIXED

People

(Reporter: bugzilla, Unassigned)

Details

I'm seeing ~60GB data on disk now, which is starting to overwhelm the single-node c3.4xlarge clusters set up for pioneer data analysis (especially since we're only at ~20GB out of 30GB free memory without spark running). Given this study will go on for 2 more weeks, we're probably looking at about 200GB of data total. My personal preference would be to have all of this fit in memory, so I think we should consider spinning up a dedicated 10-node cluster or two.
I have set up two new 10-node clusters and have (finally) automated the provisioning of such. I've updated the sections in https://mana.mozilla.org/wiki/display/SVCOPS/Access+to+Pioneer+Resources#AccesstoPioneerResources-AccessingJupyterandZeppelinNotebooks and https://mana.mozilla.org/wiki/display/SVCOPS/Access+to+Pioneer+Resources#AccesstoPioneerResources-ResourceContention to reflect the addition of these new clusters and how to access them.
Status: NEW → RESOLVED
Closed: 8 years ago
Resolution: --- → FIXED
Product: Data Platform and Tools → Data Platform and Tools Graveyard
You need to log in before you can comment on or make changes to this bug.