Szarlat J1
1XENON Systems, Springvale, Australia
Biography:
Jakub is an experienced Senior Solutions Architect with broad experience in HPC, AI, data storage and management, scientific applications, software stacks and workflows.
With a degree in Software Engineering, Jakub has a unique ability of bridging distinct systems to work as a total solution.
Jakub is an expert in architecting, deploying, and managing traditional and virtualised HPC clusters, large, distributed data storage and archive solutions, and cloud computing environments for a variety of use cases ranging from Microscopy VDI to Bioinformatics and other GPU workloads. In addition, he is supporting researchers compiling, debugging, deploying, and running their applications.
His exposure to large scientific data sets brings a world of expertise in understanding their implication on today's infrastructure and data management.
More recently at XENON, Jakub has been utilising container technologies to provide more customisable and stable HPC deployments both on-prem and in the cloud.
Abstract:
Modern eResearch HPC platforms are more than schedulers and compute nodes; they also depend on identity, portals, storage, monitoring, logging, security, software environments, and repeatable operations. This talk presents XENON Cluster Stack (XCS) as a reproducible, image-based approach to HPC management that treats node environments and services as deployable artefacts rather than hand-built systems.
XCS builds login and compute images from automation, supports diskless operation, and uses torrent-based image distribution so booting nodes help distribute the image as they come online. This reduces traditional image-server bottlenecks and enables large-scale boot scenarios, including testing with around 1000 nodes in a similar time envelope to much smaller deployments.
The talk also explores how AI can assist infrastructure teams by helping draft, review, test, document, and troubleshoot Ansible roles and platform definitions. The key message is that XCS combines reproducible images, peer-distributed boot, service-stack management, and AI-assisted operations to make HPC systems easier to update, recover, and sustain.