Federated HPC: A Reference Pattern for Unified On-Premises and Cloud Research Computing

Sterzl K1

1Amazon Web Services, Brisbane, Australia

Biography:

Kurt Sterzl is a Senior Solutions Architect with AWS Worldwide Public Sector, based in Brisbane. He works with public sector and higher-education research institutions to design and build high-performance computing, cloud, and AI/ML solutions, with a focus on practical, reproducible architectures that researchers can actually adopt. His recent work centres on hybrid HPC – federating on-premises and cloud clusters under a single scheduler while keeping data governance and workflow reproducibility intact. He is a regular collaborator with university research-computing teams across Australia, helping translate emerging cloud capability into real research outcomes.

Abstract:

Research institutions run high-performance computing (HPC) across two worlds: established on-premises clusters holding large or sensitive datasets, and cloud capacity offering elasticity and specialised hardware such as graphics processing units (GPUs).  Bridging them is painful – researchers are pushed onto a second scheduler, a second copy of their data, and software that behaves differently in each environment, producing duplicated effort, data-governance risk, and workflows that are hard to reproduce.

This presentation describes a reference pattern that federates an on-premises Slurm cluster with a cloud HPC cluster under a single scheduling fabric. A researcher submits one job and chooses where it runs with two flags, while both clusters share one view of work and accounting. Data stays authoritative on-premises and is transparently cached in the cloud on demand: the on-premises store remains the single source of truth, only the working set a job reads is pulled across, and there is no manual staging or second copy to keep in sync – so the authoritative copy never leaves the institution and only the data each job needs is transferred. Workloads run as portable containers replicated to both sides, giving identical, reproducible execution. Cloud compute scales from zero, with cost attribution and observability built in.

We walk through the architecture, share what worked and what broke, and detail the operational hazards practitioners must plan for. The environment is defined as infrastructure-as-code and validated end-to-end, so attendees leave with a concrete, reproducible blueprint for unified on-premises and cloud HPC – without sacrificing governance or reproducibility.

 

Categories

Website Sponsor

Website Sponsor