University of Cambridge Computer Laboratory - GPU cluster maintenance: VM hosts and storage – Maintenance details

GPU cluster maintenance: VM hosts and storage

Completed
Scheduled for 15 September, 2026 at 16:00 – 17:23UTC

Affects

Virtual Machine Hosting

Under maintenance from 4:00 PM to 5:23 PM

GPUs

Under maintenance from 4:00 PM to 5:23 PM

Updates
  • Completed
    15 September, 2026 at 17:23UTC
    Completed
    15 September, 2026 at 17:23UTC
    Maintenance has completed successfully.
  • Update
    15 September, 2026 at 16:59UTC
    Update
    15 September, 2026 at 16:59UTC

    GPU cluster storage is now available.

    GPU VM hosting is now available (initially at reduced VM capacity; full capacity will be restored within 10 minutes).

    Shared GPU servers will be ready for use again within 15 minutes.

    Shared CPU server within 5 minutes.

    Apologies for the slight overrun to this maintenance.

  • Update
    15 September, 2026 at 16:07UTC
    Update
    15 September, 2026 at 16:07UTC

    Please do NOT attempt to start GPU VMs during this maintenance. This only slows down completion. Please check back here to find out when the maintenance is complete.

  • In progress
    15 September, 2026 at 16:00UTC
    In progress
    15 September, 2026 at 16:00UTC
    Maintenance is now in progress
  • Planned
    11 September, 2026 at 11:27UTC
    Planned
    11 September, 2026 at 11:27UTC

    The departmental GPU cluster will be undergoing maintenance in order to maintain its security and reliability.

    This will affect:

    • Personal GPU development VMs, dev-gpu-*

    • Personal CPU development VMs, dev-cpu-* (if they are running on the GPU cluster, as most are)

    • Shared GPU development servers dev-gpu-2, dev-gpu-acs

    • Shared CPU development server dev-cpu-1

    • Access to GPU cluster storage from other systems (/anfs/gpucluster, /anfs/gpudata, /anfs/gpuscratch)

    Affected personal VMs will be shut down at the start of the maintenance and cannot be started again until it is finished.

    Shared servers will be rebooted during the maintenance.

    During the maintenance we will be installing some updates on both the VM hosts (important hypervisor security improvements) and on the storage platform (NFS reliability improvements). We will also be performing a hardware upgrade on the storage platform to enable future capacity expansion (power supply upgrade, required before we can add any more SSDs to the server).