NSF grant awarded. Thus, this project is a "go" as of August 2013.

Summary information for Hoffmann researchers, as of 12/11/13:

Sol and Helios will be unavailable to researchers between Monday, Jan. 6th at 9am and approximately Thursday, Jan 9th, at noon.

Detailed information for cluster migration:

1) Reduce the amount of researchers' temporary, test user data on Sol to the bare minimum during the break.

Bottom line: If there are any files on Sol which researchers cannot afford to lose, researchers must move that data to Helios (or elsewhere) before Sol is turned off.

QUESTION: What backup will Sol have during the winter break? Versioning? EZ-Backup? Both?

2) Monday, Jan. 6th at 9am: Sol AND Helios being turned off to researchers

3) Thursday, Jan 9th, at noon: Expect Sol available to researchers.

3) On a predetermined date, researchers must plan for 1-2 days of downtime for Helios (and before Sol is available).

At this time, Sol will still not be available to researchers. Thus, during this period when Helios, too, is turned off (to allow for its final set up), both Sol and Helios will be off-line to researchers.

We obviously must coordinate this cut-over date to ensure no loss of researcher's production data and to minimize the inconvenience of the downtime

4) About 1-2 days after Helios's shut down (3, above), Sol will be again be available, in a production capacity, to researchers.

The Hoffmann group is encouraged to improve their software installations in order to create a more robust environment and to improve support outcomes.

Tasks for Hoffmann group:

Tasks for ChemIT:

Thank you!  -ChemIT


Older notes:

Next steps

Draft idea

Unknowns

Tasks and estimated timing

Top Level Task Description

Effort Est.

Assignee

Planning

 

 

Discovery/ Overview mtg

1.5 hrs

 

Vet options and conduct needs analysis to match to hardware order

1-2 weeks

 

Specify exactly the systems to order within budget. Includes iterating with vendor experts.

1 week

 

Approval

0 days

 

Order & Installation

 

 

Place & Process order

1/2 week

 

Delivery, after order is placed at Cornell

~3 weeks

 

Receive order and set-up hardware in 248 Baker Lab

1 week

 

Build New Cluster

 

 

Get head node and 1st cluster node operational with OS and cluster management software

3 weeks

 

Test / Verify / Approval

1 week

 

Convert Old Cluster

 

 

Move user accounts and data; test, prep, and do

1 week

 

Move old nodes to new cluster

1 week

 

Other provisioning models and related ideas

Buy cycles, on demand

Good for irregular high-performance demands, especially if have high peaks of need and long-lasting jobs.

Host hardware at CAC rather than with ChemIT

Hosting costs at CAC is for basic: Expert initial configuration, then keep the system current, and keep the lights running. Other service charged hourly.

Per the above rate calculator, the rate for 9 nodes (1 head node + 8 compute nodes) would be $8,291/yr. Or, $24,873 for 3 years for this service.

At current ChemIT rates, 9 nodes would be $321.84/yr. Or, $965.52 for 3 years of service.

Table, related to our options

                          Option ==>
Consideration, below:

ChemIT

CAC:
RedCloud

CAC:
Hosting

Amazon (EC3?) or
Google (Compute?)

Other ideas?

Hardware costs

$25K

-

$25K

-

 

Hardware support

Yes.

-

Yes.

-

 

OS install and configuration

Yes. CentOS 6.4

 

Yes. CentOS 6.4

 

 

Cluster and queuing management

Yes. Warewulf, with options

-

Yes. ROCKS, no options.

-

 

Research software install and configuration

Yes

No

Yes; additional cost

No

 

Application debugging and optimization support

Not usually.
Available from CAC, at additional cost?

Yes; additional cost

Yes; additional cost

No.
Available from CAC, at additional cost?