Closed
Bug 1452837
Opened 8 years ago
Closed 8 years ago
Online News v2 Data growing large for analysis cluster size
Categories
(Data Platform and Tools Graveyard :: Operations, enhancement)
Data Platform and Tools Graveyard
Operations
Tracking
(Not tracked)
RESOLVED
FIXED
People
(Reporter: bugzilla, Unassigned)
Details
I'm seeing ~60GB data on disk now, which is starting to overwhelm the single-node c3.4xlarge clusters set up for pioneer data analysis (especially since we're only at ~20GB out of 30GB free memory without spark running).
Given this study will go on for 2 more weeks, we're probably looking at about 200GB of data total. My personal preference would be to have all of this fit in memory, so I think we should consider spinning up a dedicated 10-node cluster or two.
Comment 1•8 years ago
|
||
I have set up two new 10-node clusters and have (finally) automated the provisioning of such.
I've updated the sections in https://mana.mozilla.org/wiki/display/SVCOPS/Access+to+Pioneer+Resources#AccesstoPioneerResources-AccessingJupyterandZeppelinNotebooks and https://mana.mozilla.org/wiki/display/SVCOPS/Access+to+Pioneer+Resources#AccesstoPioneerResources-ResourceContention to reflect the addition of these new clusters and how to access them.
Status: NEW → RESOLVED
Closed: 8 years ago
Resolution: --- → FIXED
Updated•3 years ago
|
Product: Data Platform and Tools → Data Platform and Tools Graveyard
You need to log in
before you can comment on or make changes to this bug.
Description
•