Key questions to Czarek, for specing new hardware, are data needs, storage to meet those needs, and subsequent backup options.Oliver's draft, 5/16/14:
================================
Subject: Can differentiating between different types of data on Matrix create better options?
Czarek,
At this point, we want to (1) explain the process and where we are within the process, and (2) obtain some information from you to help us discuss appropriate options with you.
Process:
Here are links to a couple documents showing you ChemIT's current 's :
A) A "Cluster Upgrade Process Overview|../../../../../../../../../../download/attachments/257230542/Cluster+Upgrade+Process+Overview.pdf?version=1&modificationDate=1400517355000\" which we've used for four major cluster rebuild projects in the past year. We are currently in the planning stage, and working to develop a complete project plan and schedule for this project.
B) A summary of our initial thoughts on the current 2014-2017 Matrix upgrade / expansion project|../../../../../../../../../../download/attachments/257230542/Matrix+Cluster+2014+expansion.pdf?version=1&modificationDate=1400517437000\, including a couple areas where we have questions on the best solutions. This summary lists all of the areas we think need equipment and work this year, in order to add nodes now, and throughout the project.
Information, please:
Our current main design questions are regarding storage and backup. Today, we have several great options for file storage (and therefore backup), which are separate from the head node. We would like to consider these options with you if they are possible.
The current Matrix head node's users' data seems to us in ChemIT to be a combination of (1) user data required for computational use, and (2) user data just being stored there. The latter files may include old results and scripts, as well as personal music files and other files which do not seem to benefit from Matrix's computational capabilities.
We have a couple possible suggestions on how to handle storage which may simplify the system, but the answers are based on how Matrix storage is used.
How much of the data stored on Matrix is "Static" – being saved long term, as opposed to "active" – being used to run jobs?
Some context: In most clusters in Chemistry, Cornell, and we are told, elsewhere, users are told to move their data off the cluster to other storage once their calculations are complete. This reduces the storage needed on the cluster, and improves things like backup and system recovery, and reduces the likelihood of lost data due to a problem with the cluster.
We are more than willing to continue this dialogue by email if that works for you, or we can call you in Poland and try solving these issues quickly that way.
Do you think there is a reasonable way for your users to make this distinction with their data? We ask because this is what is done on our other clusters. On those clusters, users generally only have data required for computation on the head node. And users then move to elsewhere the resulting data and the related files worth keeping. Furthermore, on all our other clusters, users are not allowed to keep non-cluster-related data, such as music files, on the cluster's storage system. Instead, those research groups invest in a dedicated, central file storage solution.
If this distinction if reasonable, we advocate separating Matrix's storage into the two types of data: (1) data for running jobs, and (2) all other data. The benefits of privileging data actually using Matrix's computational capabilities include:
1) Cheaper and more robust storage architecture to support research-specific data. For example, the total quantity of this data is much less likely to require spanning across multiple disks.
2) Cheaper and much faster backups and restores of research-specific data. This is because it's a small fraction of the total amount of data currently being stored.
3) More useful quotas on user data, preventing digital hoarding funded by Matrix. Meaningfully make usage distinctions between data actively required for current Matrix computational work and all other data being stored (but not currently, or ever, required for computational work).
If you think there is any value in making this data distinction, ChemIT has compiled information on technical options for your consideration, along with our first pass on pros/ cons, including cost comparisons, risk, and supportability.
Note that if we retain the current architecture (the status quo in which user data is in their home directories on a single disk array), we obviously will not be able to avail ourselves of the above three cited benefits. In particular, any storage-related work on Matrix involving the current or anticipate large quantity of user data will take inordinate amounts of time, and incur higher risk than smaller amounts of data would. Matrix users are currently using 2.6TB of storage, total. But you mentioned the need to grow that capability to 12TB. How much of that space was for data actively utilizing Matrix computational capabilities, and how much of that space was for simple file storage?
By having you thoughtfully answer this data distinction question, it will heavily inform all subsequent storage and backup decisions you have to make. Thank you for your expert consideration, -Oliver and the ChemIT team.
================================
RG Draft Letter 5/19/14
Czarek –
For the first phase of the Matrix expansion / rebuild we have some overall information and some questions for you.
(Attached/Linked) are a couple documents to bring you up to speed on where we're at.
1) A "Cluster Upgrade Process Overview\" which we've used for four major cluster rebuild projects in the past year. We are currently in the planning stage, and working to develop a complete project plan and schedule for this project.
2) A summary of our initial thoughts on the current 2014-2017 Matrix upgrade / expansion project\, including a couple areas where we have questions on the best solutions. This summary lists all of the areas we think need equipment and work this year, in order to add nodes now, and throughout the project.
Our main design questions are regarding storage and backup. Today, we have several great options for file storage (and therefore backup), which are separate from the head node. We would like to consider these options with you if they are possible.
We have a couple possible suggestions on how to handle storage which may simplify the system, but the answers are based on how Matrix storage is used.
How much of the data stored on Matrix is "Static" – being saved long term, as opposed to "active" – being used to run jobs?
- In most clusters in Chemistry, Cornell, and we are told, elsewhere, users are told to move their data off the cluster to other storage once their calculations are complete. This reduces the storage needed on the cluster, and improves things like backup and system recovery, and reduces the likelihood of lost data due to a problem with the cluster.
================================
MH draft 5/19/14 5:22pm
Czarek,
At this point, we want to (1) explain the process and where we are within the process, and (2) obtain some information from you to help us provide you with appropriate options.
Process:
Today, we are still in the planning stage of this project, and need to continue to work with you to develop a complete project plan and schedule. Our "Cluster Upgrade Process Overview" document provides a very generic outline to the planning, implementing, and finishing phases we have used successfully with four other Chemistry department groups. This overview allows us to build the more accurate schedule and plan as options are discussed and decisions are made. We hope this document helps you understand the general markers of progress on this project.
Question:
Can you explain to us how users are currently using the storage available to them on the Matrix headnode? How much of the data is for their current "active" calculations, and how much is for long term "archival"? You have indicated that you want to expand out to 12TB from the 6TB capacity limit that you currently have in Matrix storage, and we want to understand the purpose of all this storage.
Our 2014-2017 Matrix upgrade/expansion initial document goes thru in very technical details all the areas that need to be thought about. It might be a little overwhelming and isn't a required complete read. Basically, once we understand your data storage needs, we can suggest the best options for Matrix storage, how to best handle backup of it, and what headnode you purchase as a result of these storage decisions.
We are more than willing to continue this dialogue by email if that works for you, or we can call you in Poland and try solving these issues quickly that way.
Chemistry IT