Hollmann H1, Rees N1, Kay B2, Croucher J1, Robinson A1, Farrington R2, Wyborn L1
1National Computational Infrastructure, Canberra, Australia, 2AuScope Ltd, Melbourne, Australia
Biography:
Dr Hannes Hollmann holds a PhD from the University of Tasmania, where he applied seismic methods to Antarctic snow and ice. He now works as a Research Data Management Specialist at the National Computational Infrastructure (NCI), where he designs and maintains data pipelines for large-scale geophysical datasets. His work sits at the intersection of high-performance computing and research data infrastructure, with a focus on making complex, heterogeneous datasets accessible and analysis-ready at scale. Hannes brings both domain knowledge in geophysics and hands-on experience in the technical challenges of curating data for automated research workflows.
Abstract:
AuScope is investing in a High-Performance Computing and Data (HPCD) Geophysics Platform hosted at the National Computational Infrastructure (NCI). The platform contains passive seismic data mirrored from the AusPass portal; magnetotelluric data from AusLAMP and other broadband and long-period surveys; and Distributed Acoustic Sounding (DAS) time series data. Co-locating these datasets alongside NCI’s compute facility optimises a range of research workflows, from individual experiments to continent-wide analysis. Integrating geophysical datasets into an HPC-hosted research data repository raises practical challenges around data preparation for automated analysis, provenance traceability, and FAIR and CARE compliance.
Preparing datasets for large-scale analysis necessitates standardising metadata across survey types to internationally agreed-upon profiles and converting data into self-describing formats amenable to programmatic access and processing. Establishing traceable provenance across multiple processing levels requires deliberate pipeline design and per-level citation, ensuring attribution throughout the research workflow, from field collection to final data products. The use of persistent identifiers is key, linking the primary instrument output to downstream derivative data products and embedding machine-readable references to related individuals, institutions, and projects.
As NCI generally hosts derivative data products, Indigenous Data Sovereignty presents a difficult challenge. Implementing the CARE Principles retrospectively requires reaching back to Traditional Custodians from whose lands the primary data were collected and consultation on how primary and derivative data products are managed and accessed. This places obligations not just on the NCI repository, but also on the collectors of the primary data, and institutional stakeholders who are best placed to facilitate that engagement.