Cloud-based Plant Phenomics Data Lakehouse

Ranaweera T1,2, Smail S3, Hobern D1

1Australian Plant Phenomics Network, Urrbrae, Australia, 2La Trobe Institute for Sustainable Agriculture and Food, Department of Ecological Plant and Animal Sciences, School of Agriculture, Biomedicine and Environment, La Trobe University, Bundoora, Australia, 3Information Services – Digital Strategy & Engagement, La Trobe University, Bundoora, Australia

Biography:

Thilina Ranaweera is a Data Platform Architect at La Trobe University’s Australian Plant Phenomics Network node, with a PhD in biomedical informatics and extensive experience in software, data management, and research systems. His work focuses on designing cloud-based Data Lakehouse capabilities that transform operational plant phenomics data into governed, analytics-ready resources for reporting, collaboration, and AI/ML enabled research. His research interests include FAIR data management, semantic interoperability, biomedical and life science informatics, and scalable digital infrastructure for controlled-environment agriculture and plant phenomics research.

Abstract:

Plant phenomics research generates large-scale, heterogeneous datasets from controlled-environment facilities, imaging systems, sensors, operational databases, and manual observations. While these datasets have value for precision agriculture, controlled-environment agriculture, and data-driven discovery, they are often distributed across secure institutional storage, vendor-specific formats, local databases, and disconnected research workflows. This fragmentation limits discovery, reuse, reproducibility, and the timely application of analytics, artificial intelligence, and machine learning.

At the La Trobe University node of the Australian Plant Phenomics Network, the task was to investigate how existing cloud platform capabilities could transform secure but siloed research data assets into a governed, scalable, and analytics-ready environment. The design process focused on data flows, repeated manual effort, and a practical architecture for ingestion, curation, governance, visualisation, and analytical reuse.

A cloud-based Data Lakehouse was developed as a Minimum Viable Product using a Medallion architecture to organise plant phenomics data into raw, curated, analytical, and shareable layers. The approach applied catalogue-driven governance, role-based data provisioning, lineage, and dataset history to improve transparency across the research data lifecycle. Rather than building a bespoke platform from the ground up, the project adapted existing cloud services to meet research infrastructure needs, creating a reusable architecture pattern based on industry-standard data engineering practices.

This presentation will share the project journey, design decisions, implementation lessons, and early outcomes from approximately three terabytes of plant phenomics data, showing how cloud-enabled research infrastructure can reduce the distance between data capture and research insight from secure research data silos to a cloud Data Lakehouse.

 

 

Categories

Website Sponsor

Website Sponsor