<?xml version="1.0" encoding="UTF-8"?>
<feed xml:lang="en-US" xmlns="http://www.w3.org/2005/Atom">
  <id>tag:status.rc.fas.harvard.edu,2005:/history</id>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu"/>
  <link rel="self" type="application/atom+xml" href="https://status.rc.fas.harvard.edu/history.atom"/>
  <title>FAS Research Computing Status - Incident history</title>
  <updated>2026-10-05T13:00:00.000+00:00</updated>
  <author>
    <name>FAS Research Computing</name>
  </author>
  
<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmuflft2j006u1bp924qtnxhd</id>
  <published>2026-10-05T13:00:00.000+00:00</published>
  <updated>2026-09-24T13:54:42.304+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmuflft2j006u1bp924qtnxhd"/>
  <title> Monthly maintenance October 5th, 2026 9am-1pm.</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    
    <p><strong>Affected Components:</strong> Infiniband - Holyoke/MGHPCC, Holyoke Firewall, Network - Holyoke/MGHPCC, Holyoke-Boston fiber link (short path), Infiniband - Boston, Network - Boston, Cambridge firewall and other redundancy, Cannon Compute Cluster (Holyoke), FASRC VPN (Boston), FASSE Open OnDemand, Web Proxies, Cannon Open OnDemand, SLURM Scheduler - FASSE, seas_compute, FASSE Compute Cluster (Holyoke), Network - Cambridge, FASRC VPN (Cambridge) , Boston Compute Nodes, Login Nodes - Boston, Holyoke-Boston fiber link (long path), Login Nodes - Holyoke, SLURM Scheduler - Cannon, Kempner Cluster GPU, Kempner Cluster CPU, GPU nodes (Holyoke), FASSE login nodes</p>
    <p><small>Sep <var data-var='date'> 24</var>, <var data-var='time'>13:54:42</var> GMT+0</small><br /><strong>Identified</strong> -
  Our maintenance tasks should be completed between **9am-1pm**.

\- CDU work continues until October 2nd 8am-5pm but will not affect running jobs or new jobs.

Phase 1 is complete. See [status page](https://status.rc.fas.harvard.edu/cmtuhw9lm0ffe0wpb7auztd1t) for timeline and affected partitions.

​- Power work on row 7c October 4th 5pm - October 5th 5pm. See [status page](https://status.rc.fas.harvard.edu/cmufluldw00bi0wp9jf318iah) for affected partitions.

**NOTICES:**

* VPN: The VPN will be cut over to new hardware on Sunday Sept. 27th. connections will drop and you will need to reconnect.
* Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at &lt;https://www.rc.fas.harvard.edu/upcoming-training/&gt;
* Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at &lt;https://status.rc.fas.harvard.edu/&gt; (click Get Updates for options).

**MAINTENANCE TASKS**

Cannon cluster will be paused during this maintenance?: **YES**  
FASSE cluster will be paused during this maintenance?: **YES**

* Network maintenance  
   * Audience: Clusters (Cannon and FASSE)  
   * Impact: **The cluster will be paused during this work.**  
   Some network disruption may occur.
* Login node reboots  
   * Audience: All login nodes  
   * Impact: Login nodes will be unavailable until after maintenance
* OOD/Open OnDemand reboots  
   * Audience: All OOD users  
   * Impact: OOD will be unavailable until after maintenance
* Netscratch 90-day retention cleanup  
   * Audience; All netscratch users  
   * Update: **Due to low space on netscratch, please note that retention cleanup will be run more frequently. We are reviewing our 90 day policy with a possible age reduction.**  
   * Impact: Files older than 90 days will be removed per our [scratch policy](https://docs.rc.fas.harvard.edu/kb/policy-scratch/). _Please note that this cleanup can happen at any time, not just during maintenance._

Thank you,  
FAS Research Computing  
&lt;https://docs.rc.fas.harvard.edu/&gt;  
&lt;https://www.rc.fas.harvard.edu/upcoming-training/&gt;
  
  .</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmufluldw00bi0wp9jf318iah</id>
  <published>2026-10-04T21:00:00.000+00:00</published>
  <updated>2026-09-24T14:06:12.177+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmufluldw00bi0wp9jf318iah"/>
  <title>Power work on row 7c October 4th 5pm - October 5th 5pm </title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    
    <p><strong>Affected Components:</strong> Cannon Compute Cluster (Holyoke)</p>
    <p><small>Sep <var data-var='date'> 24</var>, <var data-var='time'>14:06:12</var> GMT+0</small><br /><strong>Identified</strong> -
  MGHPCC will be upgrading power on Pod 7c Even Side on October 4th 5pm - October 5th 5pm. This necessitates idling half the nodes on that side of the pod. A blocking reservation has been put in place to accomplish this. No jobs will be canceled but users will notice degraded scheduling throughput due to half the nodes being closed in the following partitions:

arguelles\_delgado

blackhole

conroy

davies

desai

doshi-velez

dsouza

eddy

edwards

geophysics

giribet

gpu\_test

hernquist

huce\_cascade

huttenhower

imasc

jacobsen2

janson\_cascade

janson

ke

lukin

murphy

nguyen

ni\_lab

olveczky

ortegahernandez

pehlevan

seas\_compute

shared

shakhnovich

tambe

unrestricted

vishwanath

whipple

xlin

yin

zon
  
  .</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmun2qnu101911bnx8tw4c04j</id>
  <published>2026-09-29T19:33:25.174+00:00</published>
  <updated>2026-09-29T19:33:25.174+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmun2qnu101911bnx8tw4c04j"/>
  <title>h-nfs11-p down</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    
    <p><strong>Affected Components:</strong> Holyoke Tier 2 NFS</p>
    <p><small>Sep <var data-var='date'> 29</var>, <var data-var='time'>19:33:25</var> GMT+0</small><br /><strong>Investigating</strong> -
  h-nfs11-p is currently down. The following shares are inaccessible on FASSE

* ncf\_cnl01
* ncf\_cnl02
* ncf\_cnl03
* ncf\_cnl04
* ncf\_gspdata
* ncf\_mri\_l3
* cnl05
* cnl06
* eldaief
* ncf\_apps
* jcamprodon\_lab\_l3

We are investigating this incident..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmuldj10i004i13p57pk6h88k</id>
  <published>2026-09-28T14:59:52.672+00:00</published>
  <updated>2026-09-28T15:29:52.700+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmuldj10i004i13p57pk6h88k"/>
  <title>Authentication outage</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 30 minutes</p>
    <p><strong>Affected Components:</strong> Authentication</p>
    <p><small>Sep <var data-var='date'> 28</var>, <var data-var='time'>15:29:52</var> GMT+0</small><br /><strong>Resolved</strong> -
  Openauth/radius is now operational. This update was created by an automated monitoring service..</p>
<p><small>Sep <var data-var='date'> 28</var>, <var data-var='time'>14:59:52</var> GMT+0</small><br /><strong>Investigating</strong> -
  Authentication issues with openauth/radius. This incident was created by an automated monitoring service..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmukvtuj701t307lau15jxlta</id>
  <published>2026-09-28T06:44:22.639+00:00</published>
  <updated>2026-09-28T06:44:22.639+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmukvtuj701t307lau15jxlta"/>
  <title>FASRC VPN (Boston) is back up</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    
    <p><strong>Affected Components:</strong> FASRC VPN (Boston)</p>
    <p><small>Sep <var data-var='date'> 28</var>, <var data-var='time'>06:44:22</var> GMT+0</small><br /><strong>Investigating</strong> -
  FASRC VPN (Boston) is down at the moment. This incident was automatically created by Instatus monitoring..</p>
<p><small>Sep <var data-var='date'> 28</var>, <var data-var='time'>07:03:53</var> GMT+0</small><br /><strong>Resolved</strong> -
  FASRC VPN (Boston) is back up. This incident was automatically resolved by Instatus monitoring..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmukv7ven016p07ne59zhcooi</id>
  <published>2026-09-28T06:27:18.837+00:00</published>
  <updated>2026-09-28T06:36:45.532+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmukv7ven016p07ne59zhcooi"/>
  <title>FASRC VPN (Cambridge)  is back up</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    
    <p><strong>Affected Components:</strong> FASRC VPN (Cambridge) </p>
    <p><small>Sep <var data-var='date'> 28</var>, <var data-var='time'>06:36:45</var> GMT+0</small><br /><strong>Resolved</strong> -
  FASRC VPN (Cambridge)  is back up. This incident was automatically resolved by Instatus monitoring..</p>
<p><small>Sep <var data-var='date'> 28</var>, <var data-var='time'>06:27:18</var> GMT+0</small><br /><strong>Investigating</strong> -
  FASRC VPN (Cambridge)  is down at the moment. This incident was automatically created by Instatus monitoring..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmukuh65q01cb07oazx9l0du4</id>
  <published>2026-09-28T06:06:33.144+00:00</published>
  <updated>2026-09-28T06:06:33.144+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmukuh65q01cb07oazx9l0du4"/>
  <title>FASRC VPN (Cambridge)  is back up</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    
    <p><strong>Affected Components:</strong> FASRC VPN (Cambridge) </p>
    <p><small>Sep <var data-var='date'> 28</var>, <var data-var='time'>06:06:33</var> GMT+0</small><br /><strong>Investigating</strong> -
  FASRC VPN (Cambridge)  is down at the moment. This incident was automatically created by Instatus monitoring..</p>
<p><small>Sep <var data-var='date'> 28</var>, <var data-var='time'>06:16:03</var> GMT+0</small><br /><strong>Resolved</strong> -
  FASRC VPN (Cambridge)  is back up. This incident was automatically resolved by Instatus monitoring..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmukt0tfd01es07qne6oka72q</id>
  <published>2026-09-28T05:25:50.545+00:00</published>
  <updated>2026-09-28T05:25:50.545+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmukt0tfd01es07qne6oka72q"/>
  <title>FASRC VPN (Cambridge)  is back up</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    
    <p><strong>Affected Components:</strong> FASRC VPN (Cambridge) </p>
    <p><small>Sep <var data-var='date'> 28</var>, <var data-var='time'>05:25:50</var> GMT+0</small><br /><strong>Investigating</strong> -
  FASRC VPN (Cambridge)  is down at the moment. This incident was automatically created by Instatus monitoring..</p>
<p><small>Sep <var data-var='date'> 28</var>, <var data-var='time'>05:45:20</var> GMT+0</small><br /><strong>Resolved</strong> -
  FASRC VPN (Cambridge)  is back up. This incident was automatically resolved by Instatus monitoring..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmukr7lsv01ai07p2ny6s20iw</id>
  <published>2026-09-28T04:35:07.964+00:00</published>
  <updated>2026-09-28T04:35:07.964+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmukr7lsv01ai07p2ny6s20iw"/>
  <title>FASRC VPN (Cambridge)  is back up</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    
    <p><strong>Affected Components:</strong> FASRC VPN (Cambridge) </p>
    <p><small>Sep <var data-var='date'> 28</var>, <var data-var='time'>04:35:07</var> GMT+0</small><br /><strong>Investigating</strong> -
  FASRC VPN (Cambridge)  is down at the moment. This incident was automatically created by Instatus monitoring..</p>
<p><small>Sep <var data-var='date'> 28</var>, <var data-var='time'>04:44:38</var> GMT+0</small><br /><strong>Resolved</strong> -
  FASRC VPN (Cambridge)  is back up. This incident was automatically resolved by Instatus monitoring..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmtuhw9lm0ffe0wpb7auztd1t</id>
  <published>2026-09-22T11:00:00.000+00:00</published>
  <updated>2026-09-22T11:00:01.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmtuhw9lm0ffe0wpb7auztd1t"/>
  <title>Cooling Distribution Unit (CDU) work 9/22 - 10/2 See details for affected partitions and dates</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    
    <p><strong>Affected Components:</strong> Cannon Compute Cluster (Holyoke), SLURM Scheduler - FASSE, seas_compute, FASSE Compute Cluster (Holyoke), Boston Compute Nodes, SLURM Scheduler - Cannon, Kempner Cluster GPU, Kempner Cluster CPU, GPU nodes (Holyoke)</p>
    <p><small>Sep <var data-var='date'> 22</var>, <var data-var='time'>11:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Sep <var data-var='date'> 24</var>, <var data-var='time'>17:32:44</var> GMT+0</small><br /><strong>Identified</strong> -
  CDU 5 work is complete. 

CDU 6 work will commence..</p>
<p><small>Sep <var data-var='date'> 9</var>, <var data-var='time'>19:32:22</var> GMT+0</small><br /><strong>Identified</strong> -
  Over the past year we have noticed that water circulating through two of the Cooling Distribution Units (CDU) in Row 8a and their compute nodes has become significantly discoloured. This impurity causes cooling problems which as a result increases chances of node failure. To remedy FASRC has scheduled a full flush of these CDUs and their attached nodes. Unfortunately to do the full flush we have to fully power down all the nodes and drain the water, which is a process that takes several days to complete.

To minimize disruption we have staggered this work over two weeks. The work on CDU5 will take place from 9/22 - 9/25, while CDU6 will take place from 9/29 - 10/2\. Lists of impacted partitions are below. For those impacted we recommend using other resources on the cluster during that time such as shared and gpu\_h200 or any of the requeue partitions. Note this work impacts both Cannon and FASSE.

**No jobs will be cancelled, rather blocking reservations are in place to naturally drain the nodes prior to the work.**  

Thank you for your patience as we work to improve cluster stability and hardware longevity.  

# CDU 5 (9/22 - 9/25) -Complete

**~~Cannon~~**~~:~~

~~arguelles\_delgado\_gpu\_a100~~

~~arguelles\_delgado\_gpu\_mixed~~

~~bigmem\_intermediate~~

~~blackhole\_gpu~~

~~eddy~~

~~gershman~~

~~hejazi~~

~~hernquist\_ice~~

~~hoekstra~~

~~huce\_ice~~

~~iaifi\_gpu~~

~~itc\_gpu~~

~~jshapiro~~

~~kovac~~

~~kozinsky~~

~~kozinsky\_gpu~~

~~murphy\_ic~~

~~ortegahernandez\_ice~~

~~rivas~~

~~seas\_compute~~

~~seas\_gpu~~

~~siag~~

~~siag\_gpu~~

~~siag\_combo~~

~~sur~~

~~zhuang~~ 

**~~FASSE~~**~~:~~

~~fasse\_ultramem~~  

# CDU 6 (9/29 - 10/2)

**Cannon**:

arguelles\_delgado\_h100

bigmem

dvorkin

eddy

enos

gpu

hsph

hsph\_gpu

intermediate

itc\_cluster

janson\_sapphire

joonholee

jshapiro

olveczky\_sapphire

sapphire

test

yao

yao\_alphatns

yao\_gpu  

**FASSE**:

cnl.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmubao39814h80wruzoy3e36s</id>
  <published>2026-09-21T13:42:08.088+00:00</published>
  <updated>2026-09-21T13:42:08.088+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmubao39814h80wruzoy3e36s"/>
  <title>boslogin08 down</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 9 minutes</p>
    <p><strong>Affected Components:</strong> Login Nodes - Boston</p>
    <p><small>Sep <var data-var='date'> 21</var>, <var data-var='time'>13:42:08</var> GMT+0</small><br /><strong>Investigating</strong> -
  boslogin08 is stuck and needs to be rebooted. Please save any open work. 

We are currently investigating the underlying cause of login node issues. .</p>
<p><small>Sep <var data-var='date'> 21</var>, <var data-var='time'>13:51:34</var> GMT+0</small><br /><strong>Resolved</strong> -
  boslogin08 is back up. 

This incident has been resolved..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmu72n7vh02961nru2fkvd7g2</id>
  <published>2026-09-18T14:46:25.719+00:00</published>
  <updated>2026-09-18T18:17:50.286+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmu72n7vh02961nru2fkvd7g2"/>
  <title>holylogin05 reboot</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 3 hours and 31 minutes</p>
    <p><strong>Affected Components:</strong> Login Nodes - Holyoke</p>
    <p><small>Sep <var data-var='date'> 18</var>, <var data-var='time'>18:17:50</var> GMT+0</small><br /><strong>Resolved</strong> -
  holylogin05 is back up..</p>
<p><small>Sep <var data-var='date'> 18</var>, <var data-var='time'>14:46:25</var> GMT+0</small><br /><strong>Identified</strong> -
  holylogin05 is stuck and needs to be rebooted. Please save any open work.

We are currently investigating this incident..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmu710elw01ua1nrugxthz4dh</id>
  <published>2026-09-18T00:00:00.000+00:00</published>
  <updated>2026-09-18T00:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmu710elw01ua1nrugxthz4dh"/>
  <title>holylfs06 degraded</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    
    <p><strong>Affected Components:</strong> HolyLFS06 (Tier 0)</p>
    <p><small>Sep <var data-var='date'> 18</var>, <var data-var='time'>00:00:00</var> GMT+0</small><br /><strong>Investigating</strong> -
  holylfs06 is in a very degraded or stuck state.

Expect degraded to no access until the system is back up.
  
  .</p>
<p><small>Sep <var data-var='date'> 18</var>, <var data-var='time'>18:15:04</var> GMT+0</small><br /><strong>Monitoring</strong> -
  holylfs06 has been experiencing ongoing performance and latency issues. This has impacted workflows and made it difficult for jobs to effectively complete. Our engineers have been troubleshooting the underlying causes, and identified that the general load on holylfs06 is greater than what the filesystem can process and deliver. It is not the misuse of any particular job or group, but rather the combination of all jobs and groups. In a shared filesystem, the high load ends up impacting all users and leading to a unresponsive server. 

A more technical explanation: Jobs are reading very large volumes of data, keeping the storage arrays continuously busy with big requests. The ldiskfs journal has to make a small synchronous write every time an object is created or destroyed, and it sits on the same disks as the data. A journal bound operation that should take \~1ms is taking \~1s, with some taking over 12s. OSTs can&#039;t open new transactions until the journal commits, so service threads pile up meaning hundreds are stuck in uninterruptible wait, some for 15+ minutes. The node comes up as unresponsive, and times out. This results in the slow/stuck filesystem that many of you have experienced. 

In order to mitigate this, we have been resetting holylfs06 servers and rebooting when possible, but this is not a sustainable solution. **We are asking all groups, particularly those with large sequential reads and small-file workloads to shift their workflow to netscratch if possible.** 

A visual flowchart of an optimal workflow is depicted here: &lt;https://docs.rc.fas.harvard.edu/kb/data-storage-workflow-rdm/#Data%5FStorage%5FWorkflow&gt; 

We also have recommendations on job efficiency and best practices to be kinder to fileystems here: &lt;https://docs.rc.fas.harvard.edu/kb/job-efficiency-and-optimization-best-practices/&gt; 

A long term solution will be the upcoming Compute Storage which uses a different hardware (nvme) than holylfs06 (Lustre). We expect to migrate your holylfs06 data to Compute Storage in the coming months, and will send out additional communication at that time. 

Please reach out to us at [rchelp@rc.fas.harvard.edu](mailto:rchelp@rc.fas.harvard.edu) if your group needs additional help adjusting your workflow. 

Thank you again for your understanding..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmu5sn7ih04qg0wpcbs8yjyss</id>
  <published>2026-09-17T17:18:44.390+00:00</published>
  <updated>2026-09-17T17:18:44.390+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmu5sn7ih04qg0wpcbs8yjyss"/>
  <title>holylogin07 reboot</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 1 hour and 5 minutes</p>
    <p><strong>Affected Components:</strong> Login Nodes - Holyoke</p>
    <p><small>Sep <var data-var='date'> 17</var>, <var data-var='time'>17:18:44</var> GMT+0</small><br /><strong>Investigating</strong> -
  Due to degraded performance, holylogin07 will be rebooted at 2:15pm EST. Please save any open work. 

We are currently investigating this incident..</p>
<p><small>Sep <var data-var='date'> 17</var>, <var data-var='time'>18:23:54</var> GMT+0</small><br /><strong>Resolved</strong> -
  holylogin07 has been rebooted and is accepting new logins

We are continuing to work on a fix for this incident..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmu5k91tl02h713pcxmsrvrhm</id>
  <published>2026-09-17T13:23:44.742+00:00</published>
  <updated>2026-09-17T13:23:44.742+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmu5k91tl02h713pcxmsrvrhm"/>
  <title>holylfs06 degraded - details inside</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 2 hours and 19 minutes</p>
    <p><strong>Affected Components:</strong> HolyLFS06 (Tier 0), HolyLFS05 (Tier 0)</p>
    <p><small>Sep <var data-var='date'> 17</var>, <var data-var='time'>13:23:44</var> GMT+0</small><br /><strong>Investigating</strong> -
  Some object stores on holylfs06 are in a very degraded or stuck state. We will be restarting the entire system.

Expect degraded to no access until the system is back up..</p>
<p><small>Sep <var data-var='date'> 17</var>, <var data-var='time'>15:42:54</var> GMT+0</small><br /><strong>Resolved</strong> -
  The system has returned to service..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmu33bi0a02zt13ti0jjep2bo</id>
  <published>2026-09-15T19:54:13.267+00:00</published>
  <updated>2026-09-15T19:54:13.267+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmu33bi0a02zt13ti0jjep2bo"/>
  <title>tmux/screen not working on login nodes</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 20 hours and 27 minutes</p>
    <p><strong>Affected Components:</strong> Login Nodes - Boston, Login Nodes - Holyoke, FASSE login nodes</p>
    <p><small>Sep <var data-var='date'> 15</var>, <var data-var='time'>19:54:13</var> GMT+0</small><br /><strong>Investigating</strong> -
  tmux and screen are failing on login nodes after Monday&#039;s maintenance. 

We are currently investigating this incident. .</p>
<p><small>Sep <var data-var='date'> 16</var>, <var data-var='time'>16:21:01</var> GMT+0</small><br /><strong>Resolved</strong> -
  This issue is resolved. tmux and screen are again working normally on the login nodes..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmu334tzv02b61mti75ij3s1p</id>
  <published>2026-09-15T19:49:02.247+00:00</published>
  <updated>2026-09-16T17:02:39.160+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmu334tzv02b61mti75ij3s1p"/>
  <title>Jupyter Notebooks issue</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 21 hours and 14 minutes</p>
    <p><strong>Affected Components:</strong> FASSE Open OnDemand, Software &amp; Modules, Cannon Open OnDemand</p>
    <p><small>Sep <var data-var='date'> 16</var>, <var data-var='time'>17:02:39</var> GMT+0</small><br /><strong>Resolved</strong> -
  JupyterLab 4.5.0 is now being used for new OOD Jupyter sessions. 

This incident has been resolved..</p>
<p><small>Sep <var data-var='date'> 15</var>, <var data-var='time'>19:49:02</var> GMT+0</small><br /><strong>Monitoring</strong> -
  A bug in Jupyter AI 3.1.3 is causing idle notebooks to lose data. To remedy this, we will be reverting from Jupyter Lab 4.6.3 (with Jupyter AI 3.1.3) to the previous JupyterLab 4.5.0 release (without Jupyter AI) today   
  
If you have experienced this, check auto-saved Jupyter notebook snapshots in   
`.ipynb_checkpoints/ `  
in your working directory.   
  
Existing or currently running Jupyter notebooks will not be affected by this..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmthlo2970bfg13s2fi1zlxa0</id>
  <published>2026-09-14T13:00:00.000+00:00</published>
  <updated>2026-08-31T18:56:57.470+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmthlo2970bfg13s2fi1zlxa0"/>
  <title>FASRC monthly maintenance will take place on September 14th, 2026. </title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 4 hours</p>
    <p><strong>Affected Components:</strong> Cannon Compute Cluster (Holyoke), License Servers, Cannon Open OnDemand, SLURM Scheduler - FASSE, seas_compute, FASSE Open OnDemand, FASSE Compute Cluster (Holyoke), Netscratch (Global Scratch), Boston Compute Nodes, Login Nodes - Boston, Login Nodes - Holyoke, SLURM Scheduler - Cannon, Kempner Cluster GPU, Kempner Cluster CPU, GPU nodes (Holyoke), FASSE login nodes</p>
    <p><small>Aug <var data-var='date'> 31</var>, <var data-var='time'>18:56:57</var> GMT+0</small><br /><strong>Identified</strong> -
  Our maintenance tasks should be completed between **9am-1pm**.  
Some power work on row 7c will run **9-5** but will not affect running jobs or new jobs (see below).

**NOTICES:**

* Monday, October 12 is a university holiday (Indigenous Peoples/Columbus)
* Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at &lt;https://www.rc.fas.harvard.edu/upcoming-training/&gt;
* Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at &lt;https://status.rc.fas.harvard.edu/&gt; (click Get Updates for options).

**MAINTENANCE TASKS**

Cannon cluster will be paused during this maintenance?: **YES**  
FASSE cluster will be paused during this maintenance?: **YES**

* Slurm Upgrade to 26.05.4  
   * Audience: Cluster  
   * Impact: The cluster will be paused during this maintenance
* **New**: Enable Termination of User Processes on Logout  
   * Audience: Cluster  
   * Impact: Going forward all user processes will be terminated upon login session exit on the login nodes, excluding things running in screen and tmux. Users should leverage the cluster for non-interactive processes. This is to clean up after AI agents which tend to create orphaned processes which drag down login node performance.
* **New**: watch command cadence limit  
   * Audience: Cluster  
   * Impact:The watch command will have a minimum cadence of 60s when used on commands talking to the slurm scheduler (i.e. squeue, showq, sinfo, scontrol, sdiag, lsload). Users desiring faster polling should leverage the sacct command that talks to the slurm database. In general users should not poll the scheduler more than once every minute, ideally once every 5-10 minutes. Polling more often slows the scheduler. FASRC reserves the right to ban users who who tax the scheduler with queries. For more on cluster customs and responsibilities see: &lt;https://docs.rc.fas.harvard.edu/kb/responsibilities/&gt;
* Holyoke/MGHPCC row 7c power work - **9am-5pm (CANCELLED)**  
   * ~~Audience: Cluster~~  
   * ~~Impact: MGHPCC will be upgrading power on Pod 7c Even Side on September 14th 7am-5pm. This necessitates idling half the nodes on that side of the pod. A blocking reservation has been put in place to accomplish this. No jobs will be canceled but users will notice degraded scheduling throughput due to half the nodes being closed in the following partitions:~~  
   _~~arguelles\_delgado, blackhole, conroy, davies, desai, doshi-velez, dsouza, eddy, edwards, geophysics, giribet, gpu\_test, hernquist, huce\_cascade, huttenhower, imasc, jacobsen2, janson\_cascade, janson, ke, lukin, murphy, nguyen, ni\_lab, olveczky, ortegahernandez, pehlevan, seas\_compute, shared, shakhnovich, tambe, unrestricted, vishwanath, whipple, xlin, yin, zon~~_
* Login node reboots  
   * Audience: All login nodes  
   * Impact: Login nodes will be unavailable until after maintenance
* OOD/Open OnDemand down/reboots  
   * Audience: All OOD users  
   * Impact: OOD will be unavailable until after maintenance
* Gurobi license key update  
   * Audience: Anyone who uses Gurobi software on the clusters.  
   * Impact: Jobs running Gurobi may fail. FASRC recommends waiting until maintenance is over to run Gurobi jobs.
* Netscratch 90-day retention cleanup  
   * Audience; All netscratch users  
   * Impact: Files older than 90 days will be removed per our [scratch policy](https://docs.rc.fas.harvard.edu/kb/policy-scratch/). Please note that this cleanup can happen at any time, not just during maintenance.

Thank you,  
FAS Research Computing  
&lt;https://docs.rc.fas.harvard.edu/&gt;  
[https://www.rc.fas.harvard.edu/](https://www.rc.fas.harvard.edu/upcoming-training/).</p>
<p><small>Sep <var data-var='date'> 14</var>, <var data-var='time'>13:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Sep <var data-var='date'> 14</var>, <var data-var='time'>13:08:28</var> GMT+0</small><br /><strong>Identified</strong> -
  The Holyoke/MGHPCC row 7c power work **has been CANCELLED by the facility and will be re-scheduled.**

**The following work WILL NOT take pleace today:**

* Cancelled: MGHPCC will be upgrading power on Pod 7c Even Side on September 14th 7am-5pm. This necessitates idling half the nodes on that side of the pod. A blocking reservation has been put in place to accomplish this. No jobs will be canceled but users will notice degraded scheduling throughput due to half the nodes being closed in the following partitions:  
_arguelles\_delgado, blackhole, conroy, davies, desai, doshi-velez, dsouza, eddy, edwards, geophysics, giribet, gpu\_test, hernquist, huce\_cascade, huttenhower, imasc, jacobsen2, janson\_cascade, janson, ke, lukin, murphy, nguyen, ni\_lab, olveczky, ortegahernandez, pehlevan, seas\_compute, shared, shakhnovich, tambe, unrestricted, vishwanath, whipple, xlin, yin, zon_.</p>
<p><small>Sep <var data-var='date'> 14</var>, <var data-var='time'>17:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmtn3ixyu0ct11mqjr5qhewvl</id>
  <published>2026-09-04T15:15:42.336+00:00</published>
  <updated>2026-09-04T15:15:42.336+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmtn3ixyu0ct11mqjr5qhewvl"/>
  <title>h-nfs15-p inaccessible</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 11 days and 40 minutes</p>
    <p><strong>Affected Components:</strong> Globus Data Transfer, Holyoke Tier 2 NFS</p>
    <p><small>Sep <var data-var='date'> 4</var>, <var data-var='time'>15:15:42</var> GMT+0</small><br /><strong>Investigating</strong> -
  Storage shares on h-nfs15-p are currently unavailable. This includes:

* bellono\_lab
* debivort\_lab
* shakhnovich\_lab
* doyle\_lab
* holbrook\_lab
* kou\_lab
* koutrakis\_lab
* brennan\_lab

We are currently investigating this incident..</p>
<p><small>Sep <var data-var='date'> 4</var>, <var data-var='time'>19:21:54</var> GMT+0</small><br /><strong>Identified</strong> -
  Most shares are back up and accessible for use. 

kou\_lab remains unavailable at this time.

We are continuing to work on a fix for this incident..</p>
<p><small>Sep <var data-var='date'> 15</var>, <var data-var='time'>15:55:36</var> GMT+0</small><br /><strong>Resolved</strong> -
  This incident has been resolved..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmt8un5o5004607of5qu1dqu1</id>
  <published>2026-08-25T15:58:13.089+00:00</published>
  <updated>2026-08-25T15:58:13.089+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmt8un5o5004607of5qu1dqu1"/>
  <title>docs.rc.fas.harvard.edu is back up</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    
    <p><strong>Affected Components:</strong> docs.rc.fas.harvard.edu</p>
    <p><small>Aug <var data-var='date'> 25</var>, <var data-var='time'>15:58:13</var> GMT+0</small><br /><strong>Investigating</strong> -
  docs.rc.fas.harvard.edu is down at the moment. This incident was created automatically..</p>
<p><small>Aug <var data-var='date'> 25</var>, <var data-var='time'>16:38:14</var> GMT+0</small><br /><strong>Resolved</strong> -
  docs.rc.fas.harvard.edu is back up. This incident was resolved automatically..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmt8umhuc003p07p7029zx6xo</id>
  <published>2026-08-25T15:57:45.095+00:00</published>
  <updated>2026-08-25T15:57:45.095+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmt8umhuc003p07p7029zx6xo"/>
  <title>www.rc.fas.harvard.edu is back up</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    
    <p><strong>Affected Components:</strong> www.rc.fas.harvard.edu</p>
    <p><small>Aug <var data-var='date'> 25</var>, <var data-var='time'>15:57:45</var> GMT+0</small><br /><strong>Investigating</strong> -
  www.rc.fas.harvard.edu is down at the moment. This incident was created automatically..</p>
<p><small>Aug <var data-var='date'> 25</var>, <var data-var='time'>16:37:51</var> GMT+0</small><br /><strong>Resolved</strong> -
  www.rc.fas.harvard.edu is back up. This incident was resolved automatically..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmt8t2wuk00360lp3g578cnxc</id>
  <published>2026-08-25T15:14:32.506+00:00</published>
  <updated>2026-08-25T15:28:27.363+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmt8t2wuk00360lp3g578cnxc"/>
  <title>cannon OOD / Open OnDemand is down</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 2 hours and 44 minutes</p>
    <p><strong>Affected Components:</strong> Cannon Open OnDemand</p>
    <p><small>Aug <var data-var='date'> 25</var>, <var data-var='time'>15:28:27</var> GMT+0</small><br /><strong>Identified</strong> -
  This likely affects other virtual machines including, but possibly not limited to,   
cannon ood

cbscentral\*

ssbccentral\*

mczapps

mczbase

msprl

smms

epslic1

hptc-cal

mailman

minecat

r10k-builds02

ncfservice\*.</p>
<p><small>Aug <var data-var='date'> 25</var>, <var data-var='time'>17:58:14</var> GMT+0</small><br /><strong>Resolved</strong> -
  This incident has been resolved. All hosts are back up and online. Thanks for your patience.

Please note that the rolling OS updates are still in progress, but will not affect OOD access..</p>
<p><small>Aug <var data-var='date'> 25</var>, <var data-var='time'>15:14:32</var> GMT+0</small><br /><strong>Investigating</strong> -
  We are currently investigating this issue. Details to follow..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmsdmp5fp02hr0rppwc4nnkya</id>
  <published>2026-08-24T13:00:00.000+00:00</published>
  <updated>2026-08-27T20:06:34.347+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmsdmp5fp02hr0rppwc4nnkya"/>
  <title>Rolling OS Upgrades August 24th - 27th 2026</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 3 days, 7 hours and 7 minutes</p>
    <p><strong>Affected Components:</strong> FASSE Compute Cluster (Holyoke), Cannon Compute Cluster (Holyoke), seas_compute, SLURM Scheduler - FASSE, Boston Compute Nodes, Login Nodes - Boston, Login Nodes - Holyoke, SLURM Scheduler - Cannon, Kempner Cluster GPU, Kempner Cluster CPU, Cannon Open OnDemand, FASSE Open OnDemand, GPU nodes (Holyoke), FASSE login nodes</p>
    <p><small>Aug <var data-var='date'> 27</var>, <var data-var='time'>20:06:34</var> GMT+0</small><br /><strong>Completed</strong> -
  OS upgrade work is mostly completed. 

There are a few remaining nodes in the cluster that did not successfully receive the upgrade. 

If you see nodes that are &#039;DOWN&#039; for the following reasons, they are all nodes that need manual intervention for OS upgrade. FASRC staff are aware of them and will upgrade these nodes over the next week. 

* &quot;wrong kernel&quot;
* &quot;Not responding&quot;
* &quot;unexpected reboot&quot;
* &quot;gres/gpu count lower than configured&quot;
* &quot;/n/sw not mounted&quot;

We appreciate your continued patience..</p>
<p><small>Aug <var data-var='date'> 3</var>, <var data-var='time'>19:35:00</var> GMT+0</small><br /><strong>Identified</strong> -
  FASRC will be doing OS upgrades to the latest version of Rocky 8.10 from August 24-27th. These rolling upgrades will improve cluster security and install the latest Cuda 13.2 driver.

Upgrades will happen in stages with each stage occurring between 8a-5p each day. During that period the nodes being upgraded will be unavailable. No jobs will be canceled but jobs will stay pending if all nodes in the partition are down.

Note that during this week the cluster will be in a partially upgraded state, thus jobs that span multiple nodes may become unstable due to mismatching libraries. Users concerned about this should wait until August 28th to restart runs.

Users should plan their work accordingly.

### The upgrade schedule with impacted partitions is as follows:  

**August 24th:**

FASSE  
boslogin  
private login nodes

**August 25th: Cannon Part 1**

holylogin  
arguelles\_delgado  
blackhole  
davies\_gpu  
davies  
desai  
eddy  
holy-cow  
holy-smokes  
huce\_bigmem  
huce\_cascade  
huttenhower  
jacobsen2  
janson\_bigmem  
janson\_cascade  
janson  
ke  
lukin  
nguyen  
olveczky\_gpu  
remoteviz  
seas\_compute  
shared  
sompolinsky\_gpu  
tambe  
vishwanath  
whipple  
xlin  
xlin\_ice  
zhuang\_gpu  
zhuang

**August 26th: Cannon Part 2**  
arguelles\_delgado  
conroy  
davies  
doshi-velez  
dsouza  
edwards  
geophysics  
giribet  
gpu\_test  
hernquist  
huce\_cascadeimasc  
janson  
kempner\_h200  
kempner\_rtx  
murphy  
ni\_lab  
olveczky  
ortegahernandez  
pehlevan  
seas\_compute  
shakhnovich  
shared  
unrestricted  
xlin  
yin  
zon

**August 27th: Cannon Part 3**

arguelles\_delgado\_gpu\_a100  
arguelles\_delgado\_gpu\_mixed  
arguelles\_delgado\_h100  
bigmem\_intermediate  
bigmem  
blackhole\_gpu  
dvorkin  
eddy  
enos  
gershman  
gpu  
gpu\_h200  
hejazi  
hernquist\_ice  
hoekstra  
hsph\_gpu  
hsph  
huce\_ice  
iaifi\_gpu  
intermediate  
itc\_cluster  
itc\_gpu  
janson\_sapphire  
joonholee  
jshapiro  
kempner\_h100  
kempner\_h200  
kempner  
kempner\_interactive  
kovac  
kozinsky\_gpu  
kozinsky  
murphy\_ice  
mweber\_compute  
mweber\_gpu  
olveczky\_sapphire  
ortegahernandez\_ice  
rivas  
sapphire  
seas\_compute  
seas\_gpu  
seas\_gpu\_perf  
siag\_gpu  
siag\_combo  
siag  
sur  
test  
yao\_alphatns  
yao\_gpu  
yao  
zhuang.</p>
<p><small>Aug <var data-var='date'> 24</var>, <var data-var='time'>13:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Aug <var data-var='date'> 24</var>, <var data-var='time'>19:07:46</var> GMT+0</small><br /><strong>Identified</strong> -
  Today&#039;s portion of the upgrades are completed. Tomorrow&#039;s list (Cannon part 1) can be found in this [maintenance event](https://status.rc.fas.harvard.edu/cmsdmp5fp02hr0rppwc4nnkya).   

**August 24th:**

FASSE  
boslogin  
private login nodes.</p>
<p><small>Aug <var data-var='date'> 25</var>, <var data-var='time'>19:37:24</var> GMT+0</small><br /><strong>Identified</strong> -
  Today&#039;s work is complete for Cannon Part 1 and will resume tomorrow for Cannon Part 2 

**August 25th: Cannon Part 1**

holylogin  
arguelles\_delgado  
blackhole  
davies\_gpu  
davies  
desai  
eddy  
holy-cow  
holy-smokes  
huce\_bigmem  
huce\_cascade  
huttenhower  
jacobsen2  
janson\_bigmem  
janson\_cascade  
janson  
ke  
lukin  
nguyen  
olveczky\_gpu  
remoteviz  
seas\_compute  
shared  
sompolinsky\_gpu  
tambe  
vishwanath  
whipple  
xlin  
xlin\_ice  
zhuang\_gpu  
zhuang.</p>
<p><small>Aug <var data-var='date'> 26</var>, <var data-var='time'>20:01:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Today&#039;s work is complete. The final round, Cannon Part 3, takes place tomorrow.  
  
**August 26th: Cannon Part 2**  
arguelles\_delgado  
conroy  
davies  
doshi-velez  
dsouza  
edwards  
geophysics  
giribet  
gpu\_test  
hernquist  
huce\_cascadeimasc  
janson  
kempner\_h200  
kempner\_rtx  
murphy  
ni\_lab  
olveczky  
ortegahernandez  
pehlevan  
seas\_compute  
shakhnovich  
shared  
unrestricted  
xlin  
yin  
zon.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmsyt151c00bo0lmzmqhy77c4</id>
  <published>2026-08-18T15:15:26.546+00:00</published>
  <updated>2026-08-18T19:09:45.529+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmsyt151c00bo0lmzmqhy77c4"/>
  <title>Emergency Slurm (scheduler) patch at 2pm ETA 1-2hours. Jobs will be paused</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 3 hours and 42 minutes</p>
    <p><strong>Affected Components:</strong> FASSE Compute Cluster (Holyoke), Cannon Compute Cluster (Holyoke), seas_compute, SLURM Scheduler - FASSE, Boston Compute Nodes, SLURM Scheduler - Cannon, Kempner Cluster GPU, Kempner Cluster CPU, Cannon Open OnDemand, FASSE Open OnDemand, GPU nodes (Holyoke)</p>
    <p><small>Aug <var data-var='date'> 18</var>, <var data-var='time'>19:09:45</var> GMT+0</small><br /><strong>Resolved</strong> -
  A number of jobs appear to have terminated during the upgrade for some reason. Those jobs will re-queue or, if not, will need to be re-submitted,  
We apologize for the inconvenience..</p>
<p><small>Aug <var data-var='date'> 18</var>, <var data-var='time'>15:15:26</var> GMT+0</small><br /><strong>Investigating</strong> -
  We have been manually dealing with a Slurm memory leak bug behind the scenes and now have an official fix in the form of a Slurm upgrade (26.05.3). It is necessary to get this update in place as soon as possible.

At 2pm today we will pause the cluster. Jobs will be paused, not terminated, and the scheduler will be unavailable for querying or submitting new jobs until the update is completed. Once complete, jobs will resume.  
  
Technical information about the bug:

&lt;https://support.schedmd.com/show%5Fbug.cgi?id=25685&gt;.</p>
<p><small>Aug <var data-var='date'> 18</var>, <var data-var='time'>18:57:26</var> GMT+0</small><br /><strong>Resolved</strong> -
  The update has been applied and the scheduler is returned to service. Jobs are unpaused..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmsx98m6316141aqx0vel9cig</id>
  <published>2026-08-17T13:13:37.816+00:00</published>
  <updated>2026-08-17T13:27:11.025+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmsx98m6316141aqx0vel9cig"/>
  <title>holylfs06 degraded </title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 1 hour and 34 minutes</p>
    <p><strong>Affected Components:</strong> HolyLFS06 (Tier 0)</p>
    <p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>13:27:11</var> GMT+0</small><br /><strong>Investigating</strong> -
  holylfs06 has since become unresponsive. 

We are investigating..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>14:47:24</var> GMT+0</small><br /><strong>Resolved</strong> -
  holylfs06 is back up. Allow for some temporary slowness at first which will dissipate..</p>
<p><small>Aug <var data-var='date'> 17</var>, <var data-var='time'>13:13:37</var> GMT+0</small><br /><strong>Identified</strong> -
  An OSS on holylfs06 needs to be restarted. Performance may be affected in the meantime..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmsqgf40s01yn0rnwh663gbcf</id>
  <published>2026-08-12T19:00:14.848+00:00</published>
  <updated>2026-08-12T19:00:14.848+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmsqgf40s01yn0rnwh663gbcf"/>
  <title>h-nfs15-p inaccessible</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 20 hours and 6 minutes</p>
    <p><strong>Affected Components:</strong> Boston Tier 2 NFS</p>
    <p><small>Aug <var data-var='date'> 12</var>, <var data-var='time'>19:00:14</var> GMT+0</small><br /><strong>Investigating</strong> -
  Storage shares on h-nfs15-p may be inaccessible at this time. Affected labs include:

* bellono\_lab
* debivort\_lab
* shakhnovich\_lab
* doyle\_lab
* holbrook\_lab
* kou\_lab
* koutrakis\_lab
* brennan\_lab

We are currently investigating this incident. We will continue to update as we know more .</p>
<p><small>Aug <var data-var='date'> 12</var>, <var data-var='time'>19:54:27</var> GMT+0</small><br /><strong>Identified</strong> -
  Most shares on h-nfs15-p are back up and accessible at this time. 

kou\_lab storage is still down. 

 We are continuing to work on a fix for this incident. Updates to come. .</p>
<p><small>Aug <var data-var='date'> 13</var>, <var data-var='time'>15:05:57</var> GMT+0</small><br /><strong>Resolved</strong> -
  XFS journal recovery has completed, and kou\_lab is remounted. The lab should verify the filesystem and any data that may have been written after 11:00 AM on 8/11/26 (Tuesday).  
  
In simple terms, too many open files caused the server to become unresponsive, which triggered an automatic reset. After the reboot, XFS required an extended recovery period before the filesystem could be mounted again.  
  
This does not mean the lab caused the issue. The file descriptor limit is a system-wide kernel resource shared by all filesystems and workloads on the server, so any share or process running a large job could have contributed. kou\_lab was essentially collateral damage from machine-wide resource exhaustion, a “noisy neighbor” effect on a shared system..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmsqavxv000sm0zqi6u4xtxmh</id>
  <published>2026-08-12T16:25:20.638+00:00</published>
  <updated>2026-08-12T16:25:20.638+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmsqavxv000sm0zqi6u4xtxmh"/>
  <title>Openauth/Two-Factor issues for new users</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 2 days, 3 hours and 46 minutes</p>
    <p><strong>Affected Components:</strong> Authentication, FASRC Two-Factor (OpenAuth)</p>
    <p><small>Aug <var data-var='date'> 12</var>, <var data-var='time'>16:25:20</var> GMT+0</small><br /><strong>Identified</strong> -
  We have identified an issue which keeps new accounts from using their two-factor/openauth token for authentication.

This would also affect anyone resetting their token.

If you have a **new account** and are unable to authenticate to the cluster, FASRC VPN, or other FASRC services, this is why.  
Existing accounts are not affected.

We are working to resolve this as quickly as possible, but no ETA at this time. We will update this status as things change.

  
Thanks for your understanding and patience..</p>
<p><small>Aug <var data-var='date'> 13</var>, <var data-var='time'>17:50:43</var> GMT+0</small><br /><strong>Identified</strong> -
  If you have an existing 2FA token, please do not reset it at this time. 

We are continuing to work on a fix for this incident..</p>
<p><small>Aug <var data-var='date'> 14</var>, <var data-var='time'>20:11:29</var> GMT+0</small><br /><strong>Resolved</strong> -
  This issue is resolved and normal authentification for new accounts or accounts whose tokens have been reset/revoked can now log in.

In the event that you find your token is not working still, we recommend you revoke your token and get a new one. See the last and first sections here, repectively: &lt;https://docs.rc.fas.harvard.edu/kb/openauth/&gt;.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmsq4ua7r02kk20nzqib3gpan</id>
  <published>2026-08-12T13:36:07.239+00:00</published>
  <updated>2026-08-12T13:36:07.239+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmsq4ua7r02kk20nzqib3gpan"/>
  <title>holystore01 is wedged</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 32 minutes</p>
    <p><strong>Affected Components:</strong> Holystore01 (Tier 0)</p>
    <p><small>Aug <var data-var='date'> 12</var>, <var data-var='time'>13:36:07</var> GMT+0</small><br /><strong>Identified</strong> -
  holystore01 is wedged. We are failing over OST..</p>
<p><small>Aug <var data-var='date'> 12</var>, <var data-var='time'>14:07:58</var> GMT+0</small><br /><strong>Resolved</strong> -
  This incident has been resolved..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmsfvf69o00ba07mms0nf36ee</id>
  <published>2026-08-05T09:14:42.336+00:00</published>
  <updated>2026-08-05T09:19:23.215+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmsfvf69o00ba07mms0nf36ee"/>
  <title>Grafana Cloud (FASRC) is back up</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    
    <p><strong>Affected Components:</strong> Grafana Cloud (FASRC)</p>
    <p><small>Aug <var data-var='date'> 5</var>, <var data-var='time'>09:19:23</var> GMT+0</small><br /><strong>Resolved</strong> -
  Grafana Cloud (FASRC) is back up. This incident was automatically resolved by Instatus monitoring..</p>
<p><small>Aug <var data-var='date'> 5</var>, <var data-var='time'>09:14:42</var> GMT+0</small><br /><strong>Investigating</strong> -
  Grafana Cloud (FASRC) is down at the moment. This incident was automatically created by Instatus monitoring..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Incident/cmsexxw5p0cm40kmsddlcn049</id>
  <published>2026-08-04T16:00:00.000+00:00</published>
  <updated>2026-08-04T16:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/incident/cmsexxw5p0cm40kmsddlcn049"/>
  <title>Jobstats issue</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 1 day, 2 hours and 59 minutes</p>
    <p><strong>Affected Components:</strong> SLURM Scheduler - Cannon</p>
    <p><small>Aug <var data-var='date'> 4</var>, <var data-var='time'>16:00:00</var> GMT+0</small><br /><strong>Investigating</strong> -
  jobstats is not working due to an issue with the cgroup exporter. We are working on updating this tool, but there is no ETA at this time. .</p>
<p><small>Aug <var data-var='date'> 5</var>, <var data-var='time'>18:59:09</var> GMT+0</small><br /><strong>Resolved</strong> -
  Jobstats has been rebuilt and should be functional again for both CPU and GPU jobs 

This incident has been resolved..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmrmduopd01qg0kpnmrd8zylp</id>
  <published>2026-08-03T13:05:00.000+00:00</published>
  <updated>2026-07-15T17:57:35.792+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmrmduopd01qg0kpnmrd8zylp"/>
  <title>FASRC monthly maintenance August 3rd, 2026 9am-3pm</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 6 hours and 5 minutes</p>
    <p><strong>Affected Components:</strong> Globus Data Transfer, FASSE Compute Cluster (Holyoke), Boston Tier 2 NFS, Holyoke Tier 2 NFS, Holylabs, HolyLFS06 (Tier 0), bosECS, Boston Specialty Storage, Cannon Compute Cluster (Holyoke), seas_compute, SLURM Scheduler - FASSE, CEPH Storage Boston (Tier 2), Holyoke Specialty Storage, Isilon Storage Holyoke (Tier 1), holECS, BosLFS02 (Tier 0), Isilon Storage Boston (Tier 1), Home Directory Storage - Boston, Netscratch (Global Scratch), Tape - (Tier 3), Boston Compute Nodes, Login Nodes - Boston, Samba Cluster, HolyLFS04 (Tier 0), Holystore01 (Tier 0), FASRC Two-Factor (OpenAuth), HolyLFS05 (Tier 0), Login Nodes - Holyoke, SLURM Scheduler - Cannon, Kempner Cluster GPU, Kempner Cluster CPU, Cannon Open OnDemand, FASSE Open OnDemand, GPU nodes (Holyoke), FASSE login nodes</p>
    <p><small>Jul <var data-var='date'> 15</var>, <var data-var='time'>17:57:35</var> GMT+0</small><br /><strong>Identified</strong> -
  FASRC monthly maintenance will take place on August 3rd, 2026\. Our maintenance tasks should be completed between **9am-3pm**.

**Note**: This maintenance takes place at the same time as the [**firewall cutover \[details\]**](https://dashboard.instatus.com/fasrc/fasrc/maintenances/cmrjgmzwb01780rmp8pxkqwr0) and all running jobs **_will be canceled on the morning of August 3rd_**. Note the longer duration of this maintenance period: 9am-3pm

**NOTICES:**

* Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at &lt;https://www.rc.fas.harvard.edu/upcoming-training/&gt;
* Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at &lt;https://status.rc.fas.harvard.edu/&gt; (click Get Updates for options).

**MAINTENANCE TASKS**

Cannon cluster will be paused during this maintenance?: **YES**  
FASSE cluster will be paused during this maintenance?: **YES**

* Slurm Upgrade to 26.05.2  
   * Audience: Cluster  
   * Impact: The cluster will be paused during this maintenance
* [two-factor.rc.fas.harvard.edu](http://two-factor.rc.fas.harvard.edu) [OpenAuth](https://docs.rc.fas.harvard.edu/kb/openauth/) cut-over to new server  
   * Audience: New accounts or anyone requesting an OpenAuth token  
   * Impact: two-factor will be unavailable while moving to a new server
* Login node down/reboots  
   * Audience: All login nodes  
   * Impact: Login nodes will be unavailable until after maintenance
* OOD/Open OnDemand down/reboots  
   * Audience: All OOD users  
   * Impact: OOD will be unavailable until after maintenance
* Lab storage cutover  
   * Audience; Anyone who was contacted about this  
   * Impact: See email(s) from RDM to affected users
* Netscratch 90-day retention cleanup  
   * Audience; All netscratch users  
   * Impact: Files older than 90 days will be removed per our [scratch policy](https://docs.rc.fas.harvard.edu/kb/policy-scratch/). Please note that this cleanup can happen at any time, not just during maintenance.

Thank you,  
FAS Research Computing  
&lt;https://docs.rc.fas.harvard.edu/&gt;  
&lt;https://www.rc.fas.harvard.edu/&gt;.</p>
<p><small>Aug <var data-var='date'> 3</var>, <var data-var='time'>13:05:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Aug <var data-var='date'> 3</var>, <var data-var='time'>19:10:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmrjgmzwb01780rmp8pxkqwr0</id>
  <published>2026-08-03T13:00:00.000+00:00</published>
  <updated>2026-07-13T16:52:17.370+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmrjgmzwb01780rmp8pxkqwr0"/>
  <title>Data Center Firewall Replacement August 3rd 9am-3pm</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 6 hours</p>
    <p><strong>Affected Components:</strong> Globus Data Transfer, Infiniband - Holyoke/MGHPCC, FASSE Compute Cluster (Holyoke), Boston Tier 2 NFS, Holyoke Tier 2 NFS, Holylabs, Holyoke Firewall, Network - Holyoke/MGHPCC, Holyoke-Boston fiber link (short path), Infiniband - Boston, HolyLFS06 (Tier 0), bosECS, Boston Specialty Storage, Software &amp; Modules, Cannon Compute Cluster (Holyoke), Network - Boston, Cambridge firewall and other redundancy, License Servers, seas_compute, Web Proxies, SLURM Scheduler - FASSE, CEPH Storage Boston (Tier 2), Holyoke Specialty Storage, Isilon Storage Holyoke (Tier 1), holECS, Network - Cambridge, BosLFS02 (Tier 0), Isilon Storage Boston (Tier 1), Home Directory Storage - Boston, Netscratch (Global Scratch), Tape - (Tier 3), Boston Compute Nodes, Login Nodes - Boston, Samba Cluster, HolyLFS04 (Tier 0), Holystore01 (Tier 0), Holyoke-Boston fiber link (long path), HolyLFS05 (Tier 0), Login Nodes - Holyoke, SLURM Scheduler - Cannon, Kempner Cluster GPU, Kempner Cluster CPU, Cannon Open OnDemand, FASSE Open OnDemand, GPU nodes (Holyoke), FASSE login nodes</p>
    <p><small>Jul <var data-var='date'> 13</var>, <var data-var='time'>16:52:17</var> GMT+0</small><br /><strong>Identified</strong> -
  ### Data Center Firewall Replacement August 3rd 9am-3pm

NOTE: The following is in addition to our [regular monthly maintenance](https://status.rc.fas.harvard.edu/cmrmduopd01qg0kpnmrd8zylp) at the same time.  
  
On Monday August 3rd in addition to our monthly maintenance, the primary firewall in one of the data centers will be replaced. This will interrupt storage and home directory connectivity to the cluster and therefore the cluster must be idle for this cutover. General storage access will also be affected during this period.

This has been scheduled with networking to coincide with our monthly maintenance but will run longer. **9AM - 3PM**

**\- All jobs still running on August 3rd will be canceled -**

On the morning of August 3rd:

* Any running jobs will be canceled
* All partitions will be closed
* Login nodes will be shut down
* The Networking group will commence work

Once the firewall work is complete and tested, we will re-open the partitions and boot the login nodes. ETA 3PM

Due to the disruptive nature of this work, we want to advertise this with a longer advance notice. Normal maintenance emails will be sent per usual starting next week and will contain a condensed version of this information

Thank you,  
FAS Research Computing  
[rchelp@rc.fas.harvard.edu](mailto:rchelp@rc.fas.harvard.edu)  
&lt;https://docs.rc.fas.harvard.edu/&gt;  
&lt;https://www.rc.fas.harvard.edu/&gt;.</p>
<p><small>Aug <var data-var='date'> 3</var>, <var data-var='time'>13:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Aug <var data-var='date'> 3</var>, <var data-var='time'>19:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmquzgq6t0cjo2jqn8vofvrix</id>
  <published>2026-07-06T13:00:00.000+00:00</published>
  <updated>2026-06-26T13:45:03.139+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmquzgq6t0cjo2jqn8vofvrix"/>
  <title>FASRC monthly maintenance Monday July 6th, 2026 9am-1pm</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 4 hours</p>
    <p><strong>Affected Components:</strong> FASSE login nodes, Netscratch (Global Scratch), Login Nodes - Holyoke, Login Nodes - Boston, Cannon Open OnDemand, FASSE Open OnDemand</p>
    <p><small>Jun <var data-var='date'> 26</var>, <var data-var='time'>13:45:03</var> GMT+0</small><br /><strong>Identified</strong> -
  FASRC monthly maintenance will take place on July 6th 2026\. Our maintenance tasks should be completed between 9am-1pm.

Cannon cluster will be paused during this maintenance?: **NO**  
FASSE cluster will be paused during this maintenance?: **NO**

**NOTICES:**

* Friday July 3rd is a university holiday (independence Day observed)
* Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at &lt;https://www.rc.fas.harvard.edu/upcoming-training/&gt;
* Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at &lt;https://status.rc.fas.harvard.edu/&gt; (click Get Updates for options).
* We&#039;d love to hear success stories about your or your lab&#039;s use of FASRC. Submit your story [here](https://www.rc.fas.harvard.edu/user-stories/).

**MAINTENANCE TASKS**

* Domain controller replacement  
   * Audience: Internal  
   * Impact: None. End users should not see any impact.
* Reboot drained nodes in error state  
   * Audience: Cluster nodes with errors.  
   * Impact: These nodes will have been drained already in preparation. No impact on jobs on the day and the affected nodes will return to service in their respective partitions after the maintenance period.
* OOD/Open OnDemand reboots  
   * Audience: All OOD users, reboot of the head nodes.  
   * Impact: Running sessions will _not_ be affected.
* Login node reboots  
   * Audience; All login node users.  
   * Impact: Login nodes will reboot during the maintenance window.
* Netscratch 90-day retention cleanup  
   * Audience; All netscratch users  
   * Impact: Files older than 90 days will be removed per our [scratch policy](https://docs.rc.fas.harvard.edu/kb/policy-scratch/). Please note that this cleanup can happen at any time, not just during maintenance.

Thank you,  
FAS Research Computing  
&lt;https://docs.rc.fas.harvard.edu/&gt;  
&lt;https://www.rc.fas.harvard.edu/&gt;.</p>
<p><small>Jul <var data-var='date'> 6</var>, <var data-var='time'>13:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Jul <var data-var='date'> 6</var>, <var data-var='time'>17:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmouer4mn03r4amtc5shv8576</id>
  <published>2026-06-15T13:00:00.000+00:00</published>
  <updated>2026-06-15T13:00:01.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmouer4mn03r4amtc5shv8576"/>
  <title>2026 MGHPCC power downtime June 15-18, 2026</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 3 days, 8 hours and 16 minutes</p>
    <p><strong>Affected Components:</strong> Globus Data Transfer, Coldfront, Boston Data Center, GPU nodes (Holyoke), Infiniband - Holyoke/MGHPCC, FASSE Compute Cluster (Holyoke), SLURM Scheduler - FASSE, Boston Tier 2 NFS, Holyoke Tier 2 NFS, Holylabs, Holyoke Firewall, Network - Holyoke/MGHPCC, Holyoke-Boston fiber link (short path), Infiniband - Boston, HolyLFS06 (Tier 0), bosECS, Boston Specialty Storage, Software &amp; Modules, Holyoke/MGHPCC Data Center, Cannon Compute Cluster (Holyoke), Network - Boston, Cambridge firewall and other redundancy, License Servers, seas_compute, FASSE login nodes, Starfish, Web Proxies, Authentication, Virtual Infrastructure - Holyoke, FIINE billing portal, CEPH Storage Boston (Tier 2), NESE (NorthEast Storage Exchange), Isilon Storage Holyoke (Tier 1), holECS, Holyoke Specialty Storage, FASRC Downloads Site, Citrix, HolyLFS05 (Tier 0), Virtual Infrastructure - Boston, Network - Cambridge, FASRC VPN (Cambridge) , FASRC VPN (Boston), BosLFS02 (Tier 0), Isilon Storage Boston (Tier 1), Home Directory Storage - Boston, Netscratch (Global Scratch), Tape - (Tier 3), Boston Compute Nodes, portal.rc.fas.harvard.edu, Login Nodes - Holyoke, Login Nodes - Boston, Samba Cluster, HolyLFS04 (Tier 0), Holystore01 (Tier 0), FASRC Two-Factor (OpenAuth), Holyoke-Boston fiber link (long path), Grafana Cloud (FASRC), SLURM Scheduler - Cannon, Kempner Cluster GPU, Kempner Cluster CPU, Cannon Open OnDemand, FASSE Open OnDemand</p>
    <p><small>Jun <var data-var='date'> 15</var>, <var data-var='time'>13:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Jun <var data-var='date'> 15</var>, <var data-var='time'>13:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  The yearly power downtime at our Holyoke data center, MGHPCC, has been scheduled by the facility. This year&#039;s power downtime will take place on Tuesday June 15th - 18th, 2025\. There will be no June monthly maintenance as a result.

Since the facility will be powered down for two days this year, we will not be performing the usual maintenance tasks.   
That said, networking and other key infrastructure will be doing maintenance.

**IMPORTANT NOTE**: FASRC storage at both Holyoke and Boston **will be** affected and should not be expected to be available throughout the downtime. Please plan ahead accordingly.

* **Monday June 15th** \- Power-down begins at 9AM
* **Tuesday June 16th** \- Power out at MGHPCC
* **Wednesday June 17th** \- Power out at MGHPCC
* **Thursday June 18th** \- Expected return to full service by 5PM
* **Friday June 19th** \- Please note that June 19th is a university holiday

![Monday June 15th -  Power-down begins at 9AM
Tuesday June 16th - Power out at MGHPCC
Wednesday June 17th - Power out at MGHPCC
Thursday June 18th - Expected return to full service by 5PM](https://www.rc.fas.harvard.edu/wp-content/uploads/2026/05/mghpcc_powerdown_2026.jpg)

**For more detailed information and follow-up, please see:**   
&lt;https://www.rc.fas.harvard.edu/mghpcc-yearly-shutdown&gt; **or this** [**Status Page**](https://status.rc.fas.harvard.edu/).</p>
<p><small>Jun <var data-var='date'> 18</var>, <var data-var='time'>12:16:49</var> GMT+0</small><br /><strong>Identified</strong> -
  MGHPCC has completed their maintenance and restored power to the facility. 

FASRC will now begin the power-up process. Please be aware that this takes several hours.

We will update this status once complete.

NOTE: A reminder that tomorrow (Friday) is a university holiday. .</p>
<p><small>Jun <var data-var='date'> 18</var>, <var data-var='time'>20:45:50</var> GMT+0</small><br /><strong>Identified</strong> -
  Power-up is nearly complete, but a delay earlier in the day has us slightly behind. 

New ETA is 6PM..</p>
<p><small>Jun <var data-var='date'> 18</var>, <var data-var='time'>21:15:42</var> GMT+0</small><br /><strong>Completed</strong> -
  The yearly power downtime at our Holyoke data center, MGHPCC, has completed.

The clusters and storage are back online and login nodes and OOD nodes are now available.

If you have an issue/need help, please send a ticket to [rchelp@rc.fas.harvard.edu](mailto:rchelp@rc.fas.harvard.edu) with details.

IMPORTANT NOTE: Tomorrow, June 19th is a university holiday. FASRC staff will return Monday to address any lingering issues and any new tickets..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmoagu0310052elrw0kbo3tbu</id>
  <published>2026-05-18T11:00:00.000+00:00</published>
  <updated>2026-05-18T11:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmoagu0310052elrw0kbo3tbu"/>
  <title>MGHPCC power work - Part 2 May 18</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 3 days, 13 hours and 15 minutes</p>
    <p><strong>Affected Components:</strong> GPU nodes (Holyoke), FASSE Compute Cluster (Holyoke), SLURM Scheduler - FASSE, Cannon Compute Cluster (Holyoke), seas_compute, Kempner Cluster GPU, Kempner Cluster CPU, SLURM Scheduler - Cannon, Boston Compute Nodes</p>
    <p><small>May <var data-var='date'> 18</var>, <var data-var='time'>11:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  Rescheduled to May 18.</p>
<p><small>May <var data-var='date'> 18</var>, <var data-var='time'>11:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  Our Holyoke data center, MGHPCC, will be doing power work on Row 8A. This work, which is being completed over the course of 2 weeks, will bring online another power feed which will increase power capacity.

In order to do this work, it will require us to idle half the nodes in 8a for the duration of the week. This means all partitions in this row will be at half capacity. Existing jobs should drain naturally and no job should need to be canceled.

The impacted partitions are:

```
arguelles_delgado_h100
bigmem
bigmem_intermediate
blackhole_gpu
dvorkin
eddy
enos
gershman
gpu
gpu_h200
gpu_requeue
hejazi
hernquist_ice
hoekstra
hsph
hsph_gpu
huce_ice
iaifi_gpu_requeue
intermediate
itc_cluster
itc_gpu
janson_sapphire
joonholee
jshapiro
kempner
kempner_priority
kempner_dev
kempner_eng
kempner_h200_priority
kempner_h100
kempner_h100_priority
kempner_h100_priority2
kempner_h100_priority3
kempner_h100_priority4
kempner_interactive
kovac
kozinsky
kozinsky_gpu
kozinsky_requeue
murphy_ice
mweber_compute
mweber_gpu
olveczky_sapphire
ortegahernandez_ice
rivas
sapphire
seas_compute
seas_gpu
siag
siag_combo
test
yao
yao_priority
zhuang
```.</p>
<p><small>May <var data-var='date'> 18</var>, <var data-var='time'>11:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>May <var data-var='date'> 22</var>, <var data-var='time'>00:14:55</var> GMT+0</small><br /><strong>Completed</strong> -
  The power work has completed successfully. All nodes have been returned to normal service..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmoagr8fc0009m6hqfi4wd20x</id>
  <published>2026-05-11T11:00:00.000+00:00</published>
  <updated>2026-05-11T11:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmoagr8fc0009m6hqfi4wd20x"/>
  <title>MGHPCC power work - Part 1 May 11</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 5 days and 12 hours</p>
    <p><strong>Affected Components:</strong> GPU nodes (Holyoke), FASSE Compute Cluster (Holyoke), SLURM Scheduler - FASSE, Cannon Compute Cluster (Holyoke), seas_compute, Kempner Cluster GPU, Kempner Cluster CPU, SLURM Scheduler - Cannon, Boston Compute Nodes</p>
    <p><small>May <var data-var='date'> 11</var>, <var data-var='time'>11:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  Rescheduled to May 11.</p>
<p><small>May <var data-var='date'> 11</var>, <var data-var='time'>11:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  Our Holyoke data center, MGHPCC, will be doing power work on Row 8A. This work, which will occur this week and next week, will bring online another power feed which will increase power capacity.

In order to do this work, it will require us to idle half the nodes in 8a for the duration of the week. This means all partitions in this row will be at half capacity. Existing jobs should drain naturally and no job should need to be canceled.

The impacted partitions are:

```
arguelles_delgado_h100
bigmem
bigmem_intermediate
blackhole_gpu
dvorkin
eddy
enos
gershman
gpu
gpu_h200
gpu_requeue
hejazi
hernquist_ice
hoekstra
hsph
hsph_gpu
huce_ice
iaifi_gpu_requeue
intermediate
itc_cluster
itc_gpu
janson_sapphire
joonholee
jshapiro
kempner
kempner_priority
kempner_dev
kempner_eng
kempner_h200_priority
kempner_h100
kempner_h100_priority
kempner_h100_priority2
kempner_h100_priority3
kempner_h100_priority4
kempner_interactive
kovac
kozinsky
kozinsky_gpu
kozinsky_requeue
murphy_ice
mweber_compute
mweber_gpu
olveczky_sapphire
ortegahernandez_ice
rivas
sapphire
seas_compute
siag
siag_combo
test
yao
yao_priority
zhuang
```.</p>
<p><small>May <var data-var='date'> 16</var>, <var data-var='time'>23:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>
<p><small>May <var data-var='date'> 11</var>, <var data-var='time'>11:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmoa68qts000ceyvghagi44uw</id>
  <published>2026-05-04T13:00:00.000+00:00</published>
  <updated>2026-05-04T13:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmoa68qts000ceyvghagi44uw"/>
  <title>Monthly maintenance May 4th 2026 9am-1pm</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 4 hours</p>
    <p><strong>Affected Components:</strong> FASRC Two-Factor (OpenAuth), GPU nodes (Holyoke), FASSE Compute Cluster (Holyoke), SLURM Scheduler - FASSE, Cannon Compute Cluster (Holyoke), seas_compute, Cannon Open OnDemand, FASSE login nodes, Kempner Cluster GPU, Kempner Cluster CPU, FASSE Open OnDemand, Login Nodes - Holyoke, SLURM Scheduler - Cannon, Login Nodes - Boston, Netscratch (Global Scratch), Boston Compute Nodes</p>
    <p><small>May <var data-var='date'> 4</var>, <var data-var='time'>13:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  FASRC monthly maintenance will take place on May 4th 2026\. Our maintenance tasks should be completed between 9am-1pm.

**NOTICES:**

* Annual data center power downtime: The annual downtime at MGHPCC will take place June 15 - June 18\. This year&#039;s downtime will be one day longer. More details will be sent to all users next month.
* Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at &lt;https://www.rc.fas.harvard.edu/upcoming-training/&gt;
* Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at &lt;https://status.rc.fas.harvard.edu/&gt; (click Get Updates for options).

**MAINTENANCE TASKS**

Cannon cluster will be paused during this maintenance?: **YES**  
FASSE cluster will be paused during this maintenance?: **YES**

* Slurm 25.11.5 Upgrade  
   * Audience: All cluster users  
   * Impact: Jobs will be paused during the upgrade
* Reboot remaining stuck nodes from power outage  
   * Audience: N/A  
   * Impact: No visible impact to user
* Two-Factor/OpenAuth ([two-factor.rc.fas.harvard.edu](http://two-factor.rc.fas.harvard.edu)) replacement  
   * Audience: All account holders  
   * Impact: The server will be unavailable during maintenance. You will be unable to obtain a new or replacement OpenAuth token during this period.
* Domain controller replacement  
   * Audience: Internal  
   * Impact: End users should not see any impact
* OOD/Open OnDemand reboots  
   * Audience: All OOD users, reboot of the head nodes  
   * Impact: Running sessions will _not_ be affected
* Login node reboots  
   * Audience; All login node users  
   * Impact: Login nodes will reboot during the maintenance window
* Netscratch 90-day retention cleanup  
   * Audience; All netscratch users  
   * Impact: Files older than 90 days will be removed per our [scratch policy](https://docs.rc.fas.harvard.edu/kb/policy-scratch/). Please note that this cleanup can happen at any time, not just during maintenance.

Thank you,  
FAS Research Computing  
&lt;https://docs.rc.fas.harvard.edu/&gt;  
&lt;https://www.rc.fas.harvard.edu/&gt;.</p>
<p><small>May <var data-var='date'> 4</var>, <var data-var='time'>13:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>May <var data-var='date'> 4</var>, <var data-var='time'>14:56:02</var> GMT+0</small><br /><strong>Identified</strong> -
  The scheduler is re-opened and jobs un-paused. Other, non-impacting, work continues..</p>
<p><small>May <var data-var='date'> 4</var>, <var data-var='time'>17:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmo8wbelf004s7x6qrzk183z3</id>
  <published>2026-05-01T20:00:00.000+00:00</published>
  <updated>2026-05-01T20:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmo8wbelf004s7x6qrzk183z3"/>
  <title>Starfish maintenance</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 1 hour</p>
    <p><strong>Affected Components:</strong> Starfish</p>
    <p><small>May <var data-var='date'> 1</var>, <var data-var='time'>20:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  Starfish will be upgraded to the latest version on Friday, May 1st from 4pm-5pm. The service and dashboard will be down during this time. .</p>
<p><small>May <var data-var='date'> 1</var>, <var data-var='time'>20:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>May <var data-var='date'> 1</var>, <var data-var='time'>21:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmohlvitv03efhvfzdaadj1lz</id>
  <published>2026-04-30T12:00:00.000+00:00</published>
  <updated>2026-04-30T12:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmohlvitv03efhvfzdaadj1lz"/>
  <title>OpenOnDemand maintenance</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 2 hours</p>
    <p><strong>Affected Components:</strong> Cannon Open OnDemand, FASSE Open OnDemand</p>
    <p><small>Apr <var data-var='date'> 30</var>, <var data-var='time'>12:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  At 8am on Thursday April 30th we will be upgrading from Open OnDemand version 4.0.7 to 4.1.4 on both the Cannon and FASSE clusters. 

This is not expected to impact running jobs. 

This upgrade adds the Jobs-&gt;Project Manager menu item and fixes an issue that affected access to the Clusters-&gt;Shell Access menu item when using Firefox..</p>
<p><small>Apr <var data-var='date'> 30</var>, <var data-var='time'>12:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Apr <var data-var='date'> 30</var>, <var data-var='time'>14:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmoiqv16102yytc5y4eabtkxp</id>
  <published>2026-04-28T17:00:00.000+00:00</published>
  <updated>2026-04-28T17:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmoiqv16102yytc5y4eabtkxp"/>
  <title>Website security maintenance (www.rc and docs.rc) 4-28-26 1pm</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 17 minutes</p>
    <p><strong>Affected Components:</strong> docs.rc.fas.harvard.edu, www.rc.fas.harvard.edu</p>
    <p><small>Apr <var data-var='date'> 28</var>, <var data-var='time'>17:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  Security updates are required for [www.rc.fas.harvard.edu](http://www.rc.fas.harvard.edu) and [docs.rc.fas.harvard.edu](http://docs.rc.fas.harvard.edu)   
This work will take place today between 1pm and 2pm  
Both sites will be down for very short periods during the updates..</p>
<p><small>Apr <var data-var='date'> 28</var>, <var data-var='time'>17:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Apr <var data-var='date'> 28</var>, <var data-var='time'>17:16:58</var> GMT+0</small><br /><strong>Completed</strong> -
  Website maintenance has completed successfully..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmn6ac2960e2v140x1wt0rf80</id>
  <published>2026-04-06T13:00:00.000+00:00</published>
  <updated>2026-04-06T13:00:01.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmn6ac2960e2v140x1wt0rf80"/>
  <title>FASRC monthly maintenance April 6th 2026 9am-1pm</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 4 hours</p>
    <p><strong>Affected Components:</strong> FASRC Two-Factor (OpenAuth), , , Cannon Open OnDemand, FASSE login nodes, FASSE Open OnDemand, Login Nodes - Holyoke, Login Nodes - Boston, Netscratch (Global Scratch), , 
Login Nodes → 
OpenOnDemand/OOD → 
login.rc.fas.harvard.edu →</p>
    <p><small>Apr <var data-var='date'> 6</var>, <var data-var='time'>13:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Apr <var data-var='date'> 6</var>, <var data-var='time'>17:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>
<p><small>Apr <var data-var='date'> 6</var>, <var data-var='time'>13:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  FASRC monthly maintenance will take place on April 6th 2026\. Our maintenance tasks should be completed between 9am-1pm.

**NOTICES:**

* Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at &lt;https://www.rc.fas.harvard.edu/upcoming-training/&gt;
* Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at &lt;https://status.rc.fas.harvard.edu/&gt; (click Get Updates for options).
* We&#039;d love to hear success stories about your or your lab&#039;s use of FASRC. Submit your story [here](https://www.rc.fas.harvard.edu/user-stories/).

**MAINTENANCE TASKS**

Cannon cluster will be paused during this maintenance?: **NO**  
FASSE cluster will be paused during this maintenance?: **NO**

* [two-factor.rc.fas.harvard.edu](http://two-factor.rc.fas.harvard.edu) [OpenAuth](https://docs.rc.fas.harvard.edu/kb/openauth/) cut-over to new server  
   * Audience: New accounts or anyone requesting an OpenAuth token  
   * Impact: two-factor will be unavailable while moving to a new server
* RStudio Server (Open OnDemand)  
   * Audience: RStudio Server users on Cannon and FASSE  
   * Impact: We will be decommissioning some versions of RStudio Server so we can properly maintain all production versions. Versions to be decommissioned:  
         * R 4.1.3 (Bioconductor 3.14, RStudio 2022.02.0)  
         * R 4.1.0 (Bioconductor 3.13, RStudio 1.4.1717)  
         * R 4.0.3 (Bioconductor 3.12, Rstudio 1.3.1093)  
         * R 4.0.0 (Bioconductor 3.11, Rstudio 1.3.1093)  
   * If you use one of these versions, we recommend replacing it with the most recent version, R 4.4.2 (Bioconductor 3.20, RStudio 2024.12.0). You must reinstall previously installed libraries.
* Domain controller replacement  
   * Audience: Internal  
   * Impact: End users should not see any impact
* OOD/Open OnDemand reboots  
   * Audience: All OOD users, reboot of the head nodes  
   * Impact: Running sessions will _not_ be affected
* Login node reboots  
   * Audience; All login node users  
   * Impact: Login nodes will reboot during the maintenance window
* Netscratch 90-day retention cleanup  
   * Audience; All netscratch users  
   * Impact: Files older than 90 days will be removed per our [scratch policy](https://docs.rc.fas.harvard.edu/kb/policy-scratch/). Please note that this cleanup can happen at any time, not just during maintenance.

Thank you,  
FAS Research Computing  
&lt;https://docs.rc.fas.harvard.edu/&gt;  
&lt;https://www.rc.fas.harvard.edu/&gt;.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmls7rae90625fg9xy9b3jc5m</id>
  <published>2026-03-02T14:00:00.000+00:00</published>
  <updated>2026-03-02T14:00:01.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmls7rae90625fg9xy9b3jc5m"/>
  <title>FASRC monthly maintenance Monday March 2nd, 2026 9am-1pm</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 4 hours</p>
    <p><strong>Affected Components:</strong> Login Nodes - Holyoke, FASSE Compute Cluster (Holyoke), SLURM Scheduler - Cannon, , , GPU nodes (Holyoke), , SLURM Scheduler - FASSE, , Cannon Compute Cluster (Holyoke), , seas_compute, Cannon Open OnDemand, FASSE login nodes, Kempner Cluster CPU, Kempner Cluster GPU, FASSE Open OnDemand, Login Nodes - Boston, Netscratch (Global Scratch), , Boston Compute Nodes, 
Login Nodes → 
Cannon Cluster → 
OpenOnDemand/OOD → 
Kempner Cluster → 
FASSE Cluster → 
login.rc.fas.harvard.edu →</p>
    <p><small>Mar <var data-var='date'> 2</var>, <var data-var='time'>14:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Mar <var data-var='date'> 2</var>, <var data-var='time'>14:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  Monthly maintenance will take place on Monday March 2nd, 2026\. Our maintenance tasks should be completed between 9am-1pm.

**NOTICES:**

* Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at &lt;https://www.rc.fas.harvard.edu/upcoming-training/&gt;
* Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at &lt;https://status.rc.fas.harvard.edu/&gt; (click Get Updates for options).
* We&#039;d love to hear success stories about your or your lab&#039;s use of FASRC. Submit your story [here](https://www.rc.fas.harvard.edu/user-stories/).

**MAINTENANCE TASKS**

Cannon cluster will be paused during this maintenance?: **YES**  
FASSE cluster will be paused during this maintenance?: **YES**

* Slurm scheduler update  
   * Audience: All cluster users  
   * Impact: Jobs will be paused during maintenance
* OOD node reboots  
   * Audience; All Open OnDemand users  
   * Impact: OOD nodes will reboot during the maintenance window
* Login node reboots  
   * Audience: All login node users  
   * Impact: Login nodes will reboot during the maintenance window
* Netscratch retention purge  
   * Audience: All users of Netscratch  
   * Impact: Files older than 90 days will be removed. Please note that retention cleanup can and does run at any time, not just during the maintenance window.

Thank you,  
FAS Research Computing  
&lt;https://docs.rc.fas.harvard.edu/&gt;  
[https://www.rc.fas.harvard.edu/](https://www.rc.fas.harvard.edu/upcoming-training/).</p>
<p><small>Mar <var data-var='date'> 2</var>, <var data-var='time'>18:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmlvay24p0nlme0oeqgve2zzk</id>
  <published>2026-02-25T14:00:00.000+00:00</published>
  <updated>2026-02-25T14:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmlvay24p0nlme0oeqgve2zzk"/>
  <title>Starfish maintenance Feb 25, 2026 all day</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 1 day</p>
    <p><strong>Affected Components:</strong> Starfish</p>
    <p><small>Feb <var data-var='date'> 25</var>, <var data-var='time'>14:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  Starfish will be unavailable starting Wednesday, February 25th at 9AM until Thursday, February 26th at 9AM, for routine maintenance. The online dashboard will be inaccessible during this time..</p>
<p><small>Feb <var data-var='date'> 26</var>, <var data-var='time'>14:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>
<p><small>Feb <var data-var='date'> 25</var>, <var data-var='time'>14:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmkx2dbd201zt5svdsbm0pm92</id>
  <published>2026-02-19T13:00:00.000+00:00</published>
  <updated>2026-02-19T13:00:01.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmkx2dbd201zt5svdsbm0pm92"/>
  <title>NESE tape maintenance Feb 19th 2026</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 9 hours</p>
    <p><strong>Affected Components:</strong> NESE (NorthEast Storage Exchange)</p>
    <p><small>Feb <var data-var='date'> 19</var>, <var data-var='time'>13:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Feb <var data-var='date'> 19</var>, <var data-var='time'>13:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  From our partners at NESE. Details follow:

We are installing four new tape frames, which will bring the tape system raw storage capacity to 253 petabytes.

**Service Affected:** NESE Tape Service

**Maintenance Window:** 8:00 AM - 5:00 PM (EST)

* The tape service will be unavailable.
* All upgrade activities are expected to be completed on the same day.

NOTES:

* Monitor the MGHPCC Slack #nese channel for status updates and announcements
* Monitor &lt;https://nese.instatus.com/&gt; for real-time updates on progress

Subscribe to &lt;https://nese.instatus.com/subscribe/email&gt; for updates and announcements.</p>
<p><small>Feb <var data-var='date'> 19</var>, <var data-var='time'>22:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmlfnw5qt0xla10jy3plyxdx0</id>
  <published>2026-02-09T21:10:00.000+00:00</published>
  <updated>2026-02-09T21:10:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmlfnw5qt0xla10jy3plyxdx0"/>
  <title>Security updates needed for www.rc.fas.harvard.edu and docs.rc.fas.harvard.edu</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 8 minutes</p>
    <p><strong>Affected Components:</strong> docs.rc.fas.harvard.edu, www.rc.fas.harvard.edu</p>
    <p><small>Feb <var data-var='date'> 9</var>, <var data-var='time'>21:10:00</var> GMT+0</small><br /><strong>Identified</strong> -
  Security updates will require a brief interruption for our primary websites [www.rc.fas.harvard.edu](http://www.rc.fas.harvard.edu) and [docs.rc.fas.harvard.edu](http://docs.rc.fas.harvard.edu)

We will endeavour to keep this update as short as possible. Each site may be unavailable for a few minutes..</p>
<p><small>Feb <var data-var='date'> 9</var>, <var data-var='time'>21:10:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Feb <var data-var='date'> 9</var>, <var data-var='time'>21:17:51</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully..</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmkx2d5pn01lsytjor57r552r</id>
  <published>2026-02-09T13:00:00.000+00:00</published>
  <updated>2026-02-09T13:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmkx2d5pn01lsytjor57r552r"/>
  <title>NESE tape maintenance Feb 9th 2026</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 9 hours</p>
    <p><strong>Affected Components:</strong> NESE (NorthEast Storage Exchange)</p>
    <p><small>Feb <var data-var='date'> 9</var>, <var data-var='time'>13:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  From our partners at NESE. Details follow:

In the process of the tape front-end file caching system upgrade, we will be installing a new IBM Storage Scale System 6000\. We will provide an additional update for when the software integration and data transfer from the current IBM Elastic Storage System 5000 will be performed.

**Service Affected:** NESE Tape Service

**Maintenance Window: No Downtime expected**

NOTES:

* Monitor the MGHPCC Slack #nese channel for status updates and announcements
* Monitor &lt;https://nese.instatus.com/&gt; for real-time updates on progress
* Subscribe to &lt;https://nese.instatus.com/subscribe/email&gt; for updates and announcements.</p>
<p><small>Feb <var data-var='date'> 9</var>, <var data-var='time'>13:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Feb <var data-var='date'> 9</var>, <var data-var='time'>22:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmkvgvpg4095lzxnkkb2spait</id>
  <published>2026-02-02T14:00:00.000+00:00</published>
  <updated>2026-02-02T14:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmkvgvpg4095lzxnkkb2spait"/>
  <title>FASRC monthly maintenance Monday February 2nd, 2026 9am-1pm</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 4 hours</p>
    <p><strong>Affected Components:</strong> Cannon Open OnDemand, Login Nodes - Holyoke, , , FASSE Compute Cluster (Holyoke), GPU nodes (Holyoke), , SLURM Scheduler - FASSE, , SLURM Scheduler - Cannon, Cannon Compute Cluster (Holyoke), , FASSE Open OnDemand, seas_compute, FASSE login nodes, Kempner Cluster CPU, Kempner Cluster GPU, Login Nodes - Boston, Netscratch (Global Scratch), , Boston Compute Nodes, 
Login Nodes → 
Cannon Cluster → 
OpenOnDemand/OOD → 
Kempner Cluster → 
FASSE Cluster → 
login.rc.fas.harvard.edu →</p>
    <p><small>Feb <var data-var='date'> 2</var>, <var data-var='time'>14:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  Monthly maintenance will take place on Monday February 2nd, 2026\. Our maintenance tasks should be completed between 9am-1pm.

**NOTICES:**

* Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at &lt;https://www.rc.fas.harvard.edu/upcoming-training/&gt;
* Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at &lt;https://status.rc.fas.harvard.edu/&gt; (click Get Updates for options).
* We&#039;d love to hear success stories about your or your lab&#039;s use of FASRC. Submit your story [here](https://www.rc.fas.harvard.edu/user-stories/).

**MAINTENANCE TASKS**

Cannon cluster will be paused during this maintenance?: **YES**  
FASSE cluster will be paused during this maintenance?: **YES**

* MaxTime change  
   * Audience: Cluster users  
   * Impact: In order to improve scheduling efficiency and stability, we will be setting a maximum run time on all partitions that have MaxTime set to UNLIMITED to a MaxTime of 3 days. The unrestricted partition will be set to 365 days. Partitions that already have MaxTime set will retain their current setting. Partition owners wishing to set a different MaxTime for their partition should contact FASRC. Note that we do no guarantee uptime and so users should utilize checkpointing to save state in case of node failure.
* Slurm upgrade to 25.11.2  
   * Audience: All cluster users  
   * Impact: Jobs will be paused during maintenance
* OOD node reboots  
   * Audience; All Open OnDemand users  
   * Impact: OOD nodes will reboot during the maintenance window
* Login node reboots  
   * Audience; All login node users  
   * Impact: Login nodes will reboot during the maintenance window

Thank you,  
FAS Research Computing  
&lt;https://docs.rc.fas.harvard.edu/&gt;  
[https://www.rc.fas.harvard.edu/](https://www.rc.fas.harvard.edu/upcoming-training/).</p>
<p><small>Feb <var data-var='date'> 2</var>, <var data-var='time'>14:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Feb <var data-var='date'> 2</var>, <var data-var='time'>18:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmk1ivije002e49jtjj5n83yl</id>
  <published>2026-01-12T14:00:00.000+00:00</published>
  <updated>2026-01-12T14:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmk1ivije002e49jtjj5n83yl"/>
  <title>FASRC monthly maintenance Monday January 12th, 2026 9am-1pm</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 4 hours</p>
    <p><strong>Affected Components:</strong> , , FASSE Compute Cluster (Holyoke), SLURM Scheduler - Cannon, GPU nodes (Holyoke), , SLURM Scheduler - FASSE, , Cannon Compute Cluster (Holyoke), , Cannon Open OnDemand, FASSE Open OnDemand, seas_compute, Login Nodes - Holyoke, FASSE login nodes, Kempner Cluster CPU, Kempner Cluster GPU, Login Nodes - Boston, , Boston Compute Nodes, 
Login Nodes → 
Cannon Cluster → 
OpenOnDemand/OOD → 
Kempner Cluster → 
FASSE Cluster → 
login.rc.fas.harvard.edu →</p>
    <p><small>Jan <var data-var='date'> 12</var>, <var data-var='time'>14:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  Monthly maintenance will take place on January 12th, 2026\. Our maintenance tasks should be completed between 9am-1pm.

**NOTICES:**

* Changes to SEAS partitions, please see tasks below.
* Changes to job age priority weighting, please see tasks below.
* Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at &lt;https://status.rc.fas.harvard.edu/&gt; (click Get Updates for options).
* We&#039;d love to hear success stories about your or your lab&#039;s use of FASRC. Submit your story [here](https://www.rc.fas.harvard.edu/user-stories/).

**MAINTENANCE TASKS**

Cannon cluster will be paused during this maintenance?: **YES**  
FASSE cluster will be paused during this maintenance?:**YES**

* Slurm upgrade to 25.11.1  
   * Audience: All cluster users (Cannon and FASSE)  
   * Impact: Jobs will be paused during maintenance
* In conjunction with SEAS we will modify seas\_gpu and seas\_compute time limits  
   * Audience: SEAS users  
   * Impact:  
   seas\_gpu: will be set to 2 days maximum  
   seas\_compute: will be set to 3 days maximum  
   Existing pending jobs longer than these limits will be set to 2 day and 3 day run times depending on partition.
* Job Age Priority Weight Change  
   * Audience: Cluster users  
   * Impact: We will be adjusting the weight applied to the priority earned by jobs by virtue of their age. Currently job priority is made up of two factors, Fairshare and Job Age. The Job Age factor is currently set such that jobs gain priority over 3 days with a maximum priority equivalent to jobs with Fairshare of 0.5\. This keeps low fairshare jobs from languishing at the bottom of the queue. With the current settings though, users with low fairshare can gain significant advantage over users with higher relative fairshare. To remedy this we will be adjusting the Job Age weight to cap out at an equivalent Fairshare of 0.1\. This will still allow jobs with 0 fairshare to gain priority and thus not languish while letting fairshare govern a wider range of higher priority jobs.
* Login node reboots  
   * Audience; All login node users  
   * Impact: Login nodes will reboot during the maintenance window
* Open OnDemand (OOD) node reboots  
   * Audienc:; All OOD users  
   * Impact: OOD nodes will reboot during the maintenance window
* Netscratch retention will run  
   * Audience: All cluster netscratch users  
   * Impact: Files older than 90 days will be removed. Please note that retention cleanup can and does run at any time, not just during the maintenance window.

Thank you,  
FAS Research Computing  
&lt;https://docs.rc.fas.harvard.edu/&gt;  
[https://www.rc.fas.harvard.edu/](https://www.rc.fas.harvard.edu/upcoming-training/).</p>
<p><small>Jan <var data-var='date'> 12</var>, <var data-var='time'>14:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Jan <var data-var='date'> 12</var>, <var data-var='time'>18:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmi4yi1m400u3o9chvqic8p37</id>
  <published>2025-12-08T11:00:00.000+00:00</published>
  <updated>2025-12-08T11:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmi4yi1m400u3o9chvqic8p37"/>
  <title>Monthly Maintenance and MGHPCC Power Work - Dec. 8, 2025 6am-6pm</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 12 hours</p>
    <p><strong>Affected Components:</strong> Login Nodes - Holyoke, , GPU nodes (Holyoke), FASSE Compute Cluster (Holyoke), Isilon Storage Holyoke (Tier 1), , SLURM Scheduler - FASSE, Virtual Infrastructure - Holyoke, , Cannon Compute Cluster (Holyoke), , Cannon Open OnDemand, SLURM Scheduler - Cannon, FASSE Open OnDemand, License Servers, seas_compute, , FASSE login nodes, Kempner Cluster CPU, Kempner Cluster GPU, Login Nodes - Boston, Virtual Infrastructure - Boston, Isilon Storage Boston (Tier 1), , Boston Compute Nodes, 
Login Nodes → 
OpenOnDemand/OOD → 
Kempner Cluster → 
FASSE Cluster → 
Cannon Cluster → 
login.rc.fas.harvard.edu →</p>
    <p><small>Dec <var data-var='date'> 8</var>, <var data-var='time'>11:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  Monthly maintenance will take place on December 8th. Our maintenance tasks should be completed between 9am-1pm. However: 

_Additionally_, MGHPCC will be performing power upgrades on the odd side of Row 8A where much of our computer resides. This is the final upgrade for this row. Current estimate for this work is a 12 hour window 6am-6pm.

A list of the affected partitions is provided at the bottom of this notice. The nodes in those partitions will be drained prior to the work and will be powered down. Once the work is completed, those nodes will be returned to service. 

**Notices:**

* New FASSE partition `fasse_gpu_h200`. This partitions has 2 H200 nodes and a 3day limit. It is available now.
* 11/26 - 11/28 are university holidays (Thanksgiving). No on-site support, FASRC staff will return on 12/1.
* Training: Upcoming training from FASRC and other sources can be found on our Training Calendar. at &lt;https://www.rc.fas.harvard.edu/upcoming-training/&gt;
* Status Page: You can subscribe to our status to receive notifications of maintenance, incidents, and their resolution at &lt;https://status.rc.fas.harvard.edu/&gt; (click Get Updates for options).
* We&#039;d love to hear success stories about your or your lab&#039;s use of FASRC. Submit your story [here](https://www.rc.fas.harvard.edu/user-stories/).

**MAINTENANCE TASKS**

Cannon cluster will be paused during this maintenance?: **PARTIAL OUTAGE/YES**  
FASSE cluster will be paused during this maintenance?: **PARTIAL OUTAGE/YES**

* Power work on Row 8A odd  
   * Audience: Users of the partitions listed below  
   * Impact: These nodes and partitions will be fully or partially down all day
* OneFS (Isilon) upgrade  
   * Audience: All Isilon (Tier 1) shares  
   * Impact: Some VMs will be impacted including Cannon OOD, CBScentral, MCZapps/MCZbase, Portal, and Rclic1 (license server)
* Slurm upgrade to 25.05.5  
   * Audience: All cluster users  
   * Impact: Jobs will be paused during maintenance
* Login node reboots  
   * Audience: All login node users  
   * Impact: Login nodes will reboot during the maintenance window

**Impacted Cannon Partitions (Full or Partial Outage):**

* arguelles\_delgado\_gpu\_a100
* arguelles\_delgado\_gpu\_mixed
* bigmem\_intermediate
* blackhole\_gpu
* eddy
* gershman
* gpu\_requeue
* hejazi
* hernquist\_ice
* hoekstra
* huce\_ice
* iaifi\_gpu
* iaifi\_gpu\_priority
* iaifi\_gpu\_requeue
* itc\_gpu
* jshapiro
* kempner
* kempner\_dev
* kempner\_priority
* kempner\_h100
* kempner\_h100\_priority
* kempner\_h100\_priority2
* kempner\_h100\_priority3
* kempner\_interactive
* kempner\_requeue
* kovac
* kozinsky
* kozinsky\_gpu
* kozinsky\_priority
* kozinsky\_requeue
* murphy\_ice
* ortegahernandez\_ice
* rivas
* seas\_compute
* seas\_gpu
* serial\_requeue
* siag\_combo
* siag\_gpu
* sur
* zhuang.</p>
<p><small>Dec <var data-var='date'> 8</var>, <var data-var='time'>11:00:01</var> GMT+0</small><br /><strong>Identified</strong> -
  Maintenance is now in progress.</p>
<p><small>Dec <var data-var='date'> 8</var>, <var data-var='time'>23:00:00</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully.</p>

        ]]>
  </content>
</entry>

<entry>
  <id>tag:status.rc.fas.harvard.edu,2005:Maintenance/cmisykt3v0ayak7rmnsl4btnt</id>
  <published>2025-12-05T14:00:00.000+00:00</published>
  <updated>2025-12-05T14:00:00.000+00:00</updated>
  <link rel="alternate" type="text/html" href="https://status.rc.fas.harvard.edu/maintenance/cmisykt3v0ayak7rmnsl4btnt"/>
  <title>holylfs04 migrations</title>

  <content type="html">
  <![CDATA[
    <p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 4 days, 1 hour and 10 minutes</p>
    <p><strong>Affected Components:</strong> HolyLFS04 (Tier 0)</p>
    <p><small>Dec <var data-var='date'> 5</var>, <var data-var='time'>14:00:00</var> GMT+0</small><br /><strong>Identified</strong> -
  The holylfs04 migration to holylfs06 has begun. All holylfs04 folders will be **read-only** for the duration of the migration, from **Friday, December 5th at 9AM until end of day on Monday, December 8th.** 

All labs with holylfs04 have been informed via email; please email [rdm@rc.fas.harvard.edu](mailto:rdm@rc.fas.harvard.edu) if you have any questions..</p>
<p><small>Dec <var data-var='date'> 9</var>, <var data-var='time'>15:09:30</var> GMT+0</small><br /><strong>Completed</strong> -
  Maintenance has completed successfully..</p>

        ]]>
  </content>
</entry>

</feed>