YE Z1, Wei F2, Zhang H3, Zheng D2
1UniLab Pty Ltd, Sydney, Australia, 2The University of New South Wales , Sydney, Australia, 3The University of Sydney, Sydney, Australia
Biography:
Zhikang Ye is a Master of Computer Science student at the University of Sydney and a contributor at UniLab Pty Ltd, working on AI-enabled research infrastructure intelligence systems. Her work focuses on designing scalable data pipelines and integrating machine learning into real-world research platforms. She has contributed to the development of an auditable attribution framework that links infrastructure usage with research outputs, supporting evidence-based impact reporting. Her interests include applied AI, data engineering, and the design of transparent and trustworthy analytical systems.
Abstract:
Research infrastructure plays a critical role in enabling scientific discovery, yet its contribution to research outputs is often underreported due to inconsistent acknowledgements and reliance on manual reporting. This work presents an auditable, AI-enabled framework for linking facility usage with research publications to improve attribution accuracy and transparency.
The framework integrates internal operational data from research infrastructure platforms with external scholarly metadata sources, including ORCID and OpenAlex. It combines deterministic matching—based on persistent identifiers and temporal alignment—with probabilistic attribution using controlled large language model (LLM) extraction from acknowledgements and methods sections of publications.
To mitigate hallucination risks, LLMs are restricted to structured evidence extraction rather than attribution decisions. Extracted outputs include contribution type, supporting text snippets, confidence scores, and linkage to internal usage records, enabling risk-based human validation. A tiered model-routing strategy is applied to balance computational cost and analytical precision across varying levels of attribution complexity.
The approach is designed for independent evaluation using domain-labelled datasets, measuring precision, recall, and incremental attribution gains relative to manual processes. Early results indicate improved traceability and increased identification of previously unreported infrastructure contributions.
This work demonstrates how auditable AI pipelines can bridge the gap between operational data and research outputs, supporting more reliable reporting, funding justification, and strategic planning within research infrastructure ecosystems.