Using Structured Data to Improve AI-Based Extraction of Environmental Reports

Kneebone L1

1Kurrawongai, , Australia

Biography:

Les has worked across government, research and social service sectors developing linked data vocabularies, including the Schools Online Thesaurus (ScOT), Public Policy Taxonomy, Biosecurity Thesaurus and National Skills Taxonomy. Les has also built grey literature databases for government- and industry-led research, including the Analysis & Policy Observatory (APO) and Biosecurity Portal.

At KurrawongAI, Les provides expert advice to clients on using and creating vocabularies. He also manages reference data for the Indigenous Data Network and Biodiversity Data Repository projects. Les leads vocabulary training efforts, ensuring clients understand how to get the most out of the tools and resources we offer.

Abstract:

Large language models and document AI systems are highly effective at extracting information from complex documents, but their outputs are inherently probabilistic and frequently require post-processing, normalisation and enrichment. This paper presents an RDF/Semantic Web based extraction architecture developed for environmental chemistry reports in which structured data techniques are used not merely to store AI outputs, but to improve the overall extraction process.

Document AI outputs are transformed into RDF and knowledge graphs, providing a semantic intermediate representation with persistent identifiers and explicit relationships. Successive SPARQL transformations repair extraction errors, merge fragmented values, canonicalise entities and enrich records using domain vocabularies. Overlay graphs preserve provenance and enable corrections to be represented independently from the original extraction. The resulting knowledge graph is deployed through Apache Jena Fuseki and Prez, supporting semantic search, transparent provenance and downstream retrieval-based applications.

The approach demonstrates how symbolic knowledge representation complements statistical AI methods. Rather than replacing domain knowledge with machine learning, structured data technologies provide deterministic enrichment, explainability and governance layers that improve the quality, consistency and reusability of AI-generated information. The project illustrates the continuing value of RDF and knowledge graphs as enabling technologies for trustworthy and maintainable AI pipelines.

 

Categories

Website Sponsor

Website Sponsor