Attention: You are using an outdated browser, device or you do not have the latest version of JavaScript downloaded and so this website may not work as expected. Please download the latest software or switch device to avoid further issues.
The problem
Integrated Climate-Nature-Economic (CNE) data help capture how our climate, ecosystems, and economic activity shape, and are shaped by, one another. Today’s CNE data landscape faces a paradox: data and advanced analytical capabilities driven by AI are more accessible than ever, yet those same data remain fragmented and siloed, making seamless, accurate, and transparent data integration challenging. Inconsistent semantic interpretation — differing definitions, taxonomies, and assumptions — of data, models, and metrics limits the transparency and trust in the CNE data ecosystem. This makes it harder for decision-makers to trust, compare, and act on CNE information they increasingly depend on.
AI based on large language models (LLM) can greatly improve data discoverability and AI agents can be capable of advanced data integration. However, heterogeneous metadata and inconsistent guidance on data’s fit-for-purpose remain significant barriers to credible, consistent CNE data reuse. LLMs thrive on consistently structured data: for individual scientific and technical fields, achieving consensus on semantics needed to structure data entails intensive, multi-year processes to inclusively develop shared, consensus terminology for that field. Cancer and biomedical research, climate modeling, and device integration on the Internet of Things offer examples of the benefits of semantic interoperability and lessons for other fields working to achieve it.
Interdisciplinary, or cross-domain, semantics are a long-recognized “hard problem,” as different fields use terms inconsistently and not all fields have reached consensus on shared semantics. While cross-domain semantic interoperability is challenging, it is a key element of AI readiness for CNE data.
What are we building?
Despite the significant challenge posed by cross-domain semantics, a growing number of approaches argue for a linguistically focused, human and machine-readable solution to the problem. In partnership with the Lincoln Institute for Land Policy and Basque Centre for Climate Change, we’re building on these lessons, addressing both human and technical challenges to climate and nature data siloing through creation of a Semantic Commons for CNE data and models. This work will enable AI-supported yet human supervised integration of data and models in ways that meet the rigorous needs of finance, policy, markets, and compliance users — ensuring that subsequent reporting is more timely, transparent, and fit-for-purpose. Work includes:
Our initial focus is on the semantic interoperability of scientific data and models related to wildfire and forests. However, the approaches being developed are applicable to a wide range of data integration challenges across the CNE space and beyond it, and we will aim to extend the approach to other problem areas in future phases of the work.
Ken Bagstad, Sonia Wang (Data Foundation)
Andrew Padilla (Lincoln Institute)
Get involved (as a developer, as a domain expert, as a data/model manager)
The Semantic Commons welcomes participation from organizations and individuals committed to advancing semantic interoperability for CNE data. If you’re working on (1) the problem of semantic interoperability for interdisciplinary data, or for a key CNE domain or (2) a provider, aggregator, or user of CNE data or models, with an interest in improving their semantic interoperability, please reach out if you're interested in learning more or getting involved.