FAS Research Computing - Status Page

HolyLFS06 (Tier 0) experiencing degraded performance

Status page for the Harvard FAS Research Computing cluster and other resources.

Cluster Utilization (VPN and FASRC login required): Cannon | FASSE


Please scroll down to see details on any Incidents or maintenance notices.
Monthly maintenance occurs on the first Monday of the month (except holidays).

GETTING HELP
Documentation: https://docs.rc.fas.harvard.edu | Account Portal https://portal.rc.fas.harvard.edu
Email: rchelp@rc.fas.harvard.edu | Support Hours


The colors shown in the bars below were chosen to increase visibility for color-blind visitors.
For higher contrast, switch to light mode at the bottom of this page if the background is dark and colors are muted.

holylfs06 degraded
  • Monitoring
    UTC
    Monitoring

    holylfs06 has been experiencing ongoing performance and latency issues. This has impacted workflows and made it difficult for jobs to effectively complete. Our engineers have been troubleshooting the underlying causes, and identified that the general load on holylfs06 is greater than what the filesystem can process and deliver. It is not the misuse of any particular job or group, but rather the combination of all jobs and groups. In a shared filesystem, the high load ends up impacting all users and leading to a unresponsive server.

    A more technical explanation: Jobs are reading very large volumes of data, keeping the storage arrays continuously busy with big requests. The ldiskfs journal has to make a small synchronous write every time an object is created or destroyed, and it sits on the same disks as the data. A journal bound operation that should take ~1ms is taking ~1s, with some taking over 12s. OSTs can't open new transactions until the journal commits, so service threads pile up meaning hundreds are stuck in uninterruptible wait, some for 15+ minutes. The node comes up as unresponsive, and times out. This results in the slow/stuck filesystem that many of you have experienced.

    In order to mitigate this, we have been resetting holylfs06 servers and rebooting when possible, but this is not a sustainable solution. We are asking all groups, particularly those with large sequential reads and small-file workloads to shift their workflow to netscratch if possible.

    A visual flowchart of an optimal workflow is depicted here: https://docs.rc.fas.harvard.edu/kb/data-storage-workflow-rdm/#Data_Storage_Workflow

    We also have recommendations on job efficiency and best practices to be kinder to fileystems here: https://docs.rc.fas.harvard.edu/kb/job-efficiency-and-optimization-best-practices/

    A long term solution will be the upcoming Compute Storage which uses a different hardware (nvme) than holylfs06 (Lustre). We expect to migrate your holylfs06 data to Compute Storage in the coming months, and will send out additional communication at that time.

    Please reach out to us at rchelp@rc.fas.harvard.edu if your group needs additional help adjusting your workflow.

    Thank you again for your understanding.

  • Investigating
    UTC
    Investigating

    holylfs06 is in a very degraded or stuck state.

    Expect degraded to no access until the system is back up.



Cooling Distribution Unit (CDU) work 9/22 - 10/2 See details for affected partitions and dates
250 hoursScheduled for September 22, 2026 at 11:00 AM – October 02, 2026 at 9:00 PMUTC
  • Planned
    September 09, 2026 at 7:32 PMUTC
    Planned
    September 09, 2026 at 7:32 PMUTC

    Over the past year we have noticed that water circulating through two of the Cooling Distribution Units (CDU) in Row 8a and their compute nodes has become significantly discoloured. This impurity causes cooling problems which as a result increases chances of node failure. To remedy FASRC has scheduled a full flush of these CDUs and their attached nodes. Unfortunately to do the full flush we have to fully power down all the nodes and drain the water, which is a process that takes several days to complete.

    To minimize disruption we have staggered this work over two weeks. The work on CDU5 will take place from 9/22 - 9/25, while CDU6 will take place from 9/29 - 10/2. Lists of impacted partitions are below. For those impacted we recommend using other resources on the cluster during that time such as shared and gpu_h200 or any of the requeue partitions. Note this work impacts both Cannon and FASSE.

    No jobs will be cancelled, rather blocking reservations are in place to naturally drain the nodes prior to the work.

    Thank you for your patience as we work to improve cluster stability and hardware longevity.

    CDU 5 (9/22 - 9/25)

    Cannon:

    arguelles_delgado_gpu_a100

    arguelles_delgado_gpu_mixed

    bigmem_intermediate

    blackhole_gpu

    eddy

    gershman

    hejazi

    hernquist_ice

    hoekstra

    huce_ice

    iaifi_gpu

    itc_gpu

    jshapiro

    kovac

    kozinsky

    kozinsky_gpu

    murphy_ic

    ortegahernandez_ice

    rivas

    seas_compute

    seas_gpu

    siag

    siag_gpu

    siag_combo

    sur

    zhuang

    FASSE:

    fasse_ultramem

    CDU 6 (9/29 - 10/2)

    Cannon:

    arguelles_delgado_h100

    bigmem

    dvorkin

    eddy

    enos

    gpu

    hsph

    hsph_gpu

    intermediate

    itc_cluster

    janson_sapphire

    joonholee

    jshapiro

    olveczky_sapphire

    sapphire

    test

    yao

    yao_alphatns

    yao_gpu

    FASSE:

    cnl

Operational

SLURM Scheduler - Cannon - Operational

Cannon Compute Cluster (Holyoke) - Operational

Boston Compute Nodes - Operational

GPU nodes (Holyoke) - Operational

seas_compute - Operational

Operational

SLURM Scheduler - FASSE - Operational

FASSE Compute Cluster (Holyoke) - Operational

Operational

Kempner Cluster CPU - Operational

Kempner Cluster GPU - Operational

Operational

FASSE login nodes - Operational

Operational

Cannon Open OnDemand - Operational

FASSE Open OnDemand - Operational

Degraded performance

Netscratch (Global Scratch) - Operational

Home Directory Storage - Boston - Operational

Tape - (Tier 3) - Operational

Holylabs - Operational

Isilon Storage Holyoke (Tier 1) - Operational

Holystore01 (Tier 0) - Operational

HolyLFS04 (Tier 0) - Operational

HolyLFS05 (Tier 0) - Operational

HolyLFS06 (Tier 0) - Degraded performance

Holyoke Tier 2 NFS - Operational

Holyoke Specialty Storage - Operational

holECS - Operational

Isilon Storage Boston (Tier 1) - Operational

BosLFS02 (Tier 0) - Operational

Boston Tier 2 NFS - Operational

CEPH Storage Boston (Tier 2) - Operational

Boston Specialty Storage - Operational

bosECS - Operational

Samba Cluster - Operational

Globus Data Transfer - Operational

Recent notices

Show notice history