Jump to content

Maintenance Schedule: Difference between revisions

From HPCwiki
Dawes001 (talk | contribs)
No edit summary
 
(14 intermediate revisions by 5 users not shown)
Line 1: Line 1:
== Maintenance and Management ==
== Anunna maintenance ==
Maintenance windows allow us to safely perform essential updates and upgrades that cannot be carried out while the cluster is in use. By grouping this work into planned downtimes, we keep Anunna secure, reliable and up to date while minimising unexpected disruption.


The cluster is maintained for firmware and software updates on a six-monthly basis. Typically one downtime is scheduled in May, whilst the other is scheduled in November, in order to disrupt usage as little as possible.
Anunna is taken down for planned maintenance — firmware and software updates — roughly twice a year. Typically one downtime is scheduled in April or May and another in October in a week free of teaching, to disrupt usage as little as possible.


Any issues should be directed to the WUR IT Servicedesk:
During a downtime the cluster is unavailable. Jobs are not scheduled to run across a planned downtime, so check the announced dates when planning long jobs.
== Maintenance schedule ==
{| class="wikitable"
! Scheduled downtime
! Date
! Planned work
! Unavailable services
|-
| Autumn 2026
| 12–16 October 2026
| OS patching<br>Lustre upgrade to version 2.17
| HPC compute<br>Login nodes<br>Web applications<br>Lustre storage
|-
| Spring 2027
|
|
|
|-
| Autumn 2027
|
|
|
|}
== Announcements ==


This can be done via the mail: servicedesk.it@wur.nl
Planned maintenance is announced in advance — watch the Message of the Day on the [[Portal Overview|Apps Portal]], login nodes and we always send out an announcement in advance over e-mail.
And it can be done via the telephone: +31 317 488888
== Reporting problems ==
Please give your name and phonenumber and tell that your mail/call is about the HPC for Agrogenomics and give the company you are working for.
When you call the servicedesk, give also your email address.


== Previous Maintenance Windows ==
If something is not working — during or outside maintenance — see [[How to Get Help]].


=== Maintenance May 24th 2017 ===
== See also ==
This update moved Bright Cluster Manager to version 7.3, and SLURM to 16.08. Downtime was 8am to 8pm.
* [[How to Get Help]]
 
* [[Cluster Architecture Overview]]
=== Maintenance November 16th 2016 ===
This was a firmware update and OS update for the system. Downtime was 8am to 8pm.
 
=== Maintenance May 23rd-29th 2016 ===
This update moved the OS from Scientific Linux 6 to Scientific Linux 7, and updated Bright Cluster manager to version 7.2. A week was taken to reconstruct the entire environment from scratch, thanks to assistance from CLustervision for this expedience.
 
=== Maintenance March 23rd 2016 ===
There will be mainly firmware maintenance between 8h and 20h CET . Because network controller and storage controller firmware will be upgrades, all servers need to be rebooted and also will have network hick ups. So to prevent job and/or data corruption, the HPC will be shut down during this maintenance window. Running jobs will be killed!
 
=== Maintenance June 17th 2015 ===
There will be mainly firmware maintenance between 8h and 20h CET . Because network controller and storage controller firmware will be upgrades, all servers need to be rebooted and also will have network hick ups. So to prevent job and/or data corruption, the HPC will be shut down during this maintenance window. Running jobs will be killed!
 
=== Maintenance November 26th 2014 ===
There will be mainly firmware maintenance between 8h and 20h CET . Because network controller and storage controller firmware will be upgrades, all servers need to be rebooted and also will have network hick ups. So to prevent job and/or data corruption, the HPC will be shut down during this maintenance window. Running jobs will be killed!
 
=== Maintenance June 11th 2014 ===
There will be mainly firmware maintenance between 8h and 13h CET . Because network controller and storage controller firmware will be upgrades, all servers need to be rebooted and also will have network hick ups. So to prevent job and/or data corruption, the HPC will be shut down during this maintenance window. Running jobs will be killed!
 
=== Maintenance November 26th 2014 ===
There will be mainly firmware maintenance between 8h and 13h CET . Because network controller and storage controller firmware will be upgrades, all servers need to be rebooted and also will have network hick ups. So to prevent job and/or data corruption, the HPC will be shut down during this maintenance window. Running jobs will be killed!

Latest revision as of 08:44, 1 September 2026

Anunna maintenance

Maintenance windows allow us to safely perform essential updates and upgrades that cannot be carried out while the cluster is in use. By grouping this work into planned downtimes, we keep Anunna secure, reliable and up to date while minimising unexpected disruption.

Anunna is taken down for planned maintenance — firmware and software updates — roughly twice a year. Typically one downtime is scheduled in April or May and another in October in a week free of teaching, to disrupt usage as little as possible.

During a downtime the cluster is unavailable. Jobs are not scheduled to run across a planned downtime, so check the announced dates when planning long jobs.

Maintenance schedule

Scheduled downtime Date Planned work Unavailable services
Autumn 2026 12–16 October 2026 OS patching
Lustre upgrade to version 2.17
HPC compute
Login nodes
Web applications
Lustre storage
Spring 2027
Autumn 2027

Announcements

Planned maintenance is announced in advance — watch the Message of the Day on the Apps Portal, login nodes and we always send out an announcement in advance over e-mail.

Reporting problems

If something is not working — during or outside maintenance — see How to Get Help.

See also