Key Points

Introduction


  • Interoperability ensures that data can be understood, combined, accessed, and reused across tools, institutions, and workflows with minimal manual intervention.

  • Interoperability operates at three complementary layers:structural (how data is encoded and organized),semantic (how data is described and interpreted), and technical (how data is accessed and exchanged).

  • The FAIR interoperability principles I1–I3 primarily address the semantic layer. They provide essential guidance on shared metadata languages, vocabularies, and references, but they do not fully cover structural and technical interoperability.

  • In climate and atmospheric science, all three layers are required for practical reuse. Structural standards (e.g., NetCDF, Zarr), semantic conventions (e.g., CF), and technical mechanisms (e.g., APIs, OPeNDAP, THREDDS) must work together.

  • Many real-world barriers to reuse datasets (unclear metadata, missing units, inconsistent coordinate systems, incompatible file formats, unstable access mechanisms) are failures of one or more interoperability layers.

  • Interoperable research workflows rely on established community formats, standardized metadata conventions, stable access protocols, and scalable cloud-native layouts that allow large heterogeneous datasets to be aligned, streamed, and analysed consistently.

  • Interoperability is essential in climate science because datasets come from diverse sources (models, satellites, sensors, reanalysis) and must be combined into integrated analyses that are reproducible and machine-actionable.

Structural interoperability


  • Structural interoperability is a shared, machine-actionable contract about how data objects are organised, typed, related, and represented.
  • A file extension alone does not provide that contract: data models, encodings, schemas, and community conventions play different roles.
  • The appropriate representation depends on the underlying data model; tables, multidimensional arrays, meteorological fields, and rasters have different structural requirements.
  • CSV and TSV are highly portable but weakly self-describing; schemas and explicit structural rules make them more predictable for software.
  • NetCDF provides a shared multidimensional array data model based on dimensions, variables, and attributes.
  • NetCDF makes dataset structure machine-actionable, while conventions such as CF add community rules needed for consistent scientific interpretation.
  • Structural interoperability does not by itself guarantee semantic compatibility or technical accessibility.

Semantic interoperability


  • Semantic interoperability concerns shared, explicit, and machine-actionable meaning.
  • File formats and readable labels provide containers for meaning but do not guarantee semantic agreement.
  • Catalogues such as EarthPortal support the discovery, assessment, mapping, and sharing of Earth-science semantic artefacts, but users must still evaluate authority, versioning, provenance, licensing, and community adoption.
  • Controlled vocabulary terms, units, coordinates, bounds, cell methods, grid mappings, flags, and qualified relationships work together to express scientific meaning.
  • A CF standard_name identifies a physical quantity; long_name remains free text for human readability.
  • The same standard name and convertible units are not sufficient to guarantee direct scientific comparability.
  • CF compliance is relative to a specific version and does not prove scientific correctness, data quality, or suitability for a particular analysis.
  • The global Conventions attribute is a conformance claim, not evidence that every requirement is satisfied.
  • Compliance checkers evaluate implemented rules and must be interpreted alongside the authoritative specification and domain knowledge.
  • The IDRA files are structurally similar and well documented for human readers, but some radar semantics may still require controlled mappings, clearer unit syntax, measurement-context metadata, and provenance.
  • Semantic interoperability is achieved through community-agreed definitions, stable identifiers, explicit relationships, and transparent harmonisation—not through variable names alone.

Technical interoperability: Data access protocols


  • Technical interoperability enables machine-to-machine data exchange through standardized protocols.

  • OPeNDAP implements the DAP protocol for remote access to structured scientific datasets.

  • Remote datasets can be explored without full download.

  • Server-side subsetting reduces bandwidth and supports scalable workflows.

  • Data-access protocols can transform repositories from places where data are stored into infrastructures where data can also be accessed programmatically and selectively.

Technical interoperability: API


  • APIs operationalize technical interoperability by enabling standardized machine-to-machine interaction.

  • Web APIs use HTTP methods, predictable endpoints, JSON representations, stable identifiers, and authentication mechanisms.

  • APIs depend on structural interoperability (schemas) and semantic interoperability (controlled vocabularies).

  • Command-line tools such as curl provide direct access to API functionality and enable automation.

  • The 4TU.ResearchData API supports full dataset lifecycle management: discovery, creation, metadata update, file upload, and submission for review.

Cloud-Native Layouts


  • Cloud-native data layouts address the problem of efficiently accessing small parts of very large remote datasets.
  • Object storage manages independently addressable objects rather than providing the same access model as a traditional filesystem.
  • A conventional NetCDF file can be stored in the cloud without becoming a cloud-native data layout.
  • Zarr uses a chunk-oriented representation that maps naturally onto selective and parallel access in object storage.
  • Cloud-native layouts affect structural and technical interoperability; semantic interoperability still depends on conventions such as CF.
  • Kerchunk provides a Zarr-compatible reference view of existing NetCDF or HDF5 data without duplicating the underlying scientific data.
  • OPeNDAP, Zarr, and Kerchunk provide different approaches to efficient remote scientific-data access.

Interoperable Infrastructure in the AI Era


AI readiness is task-dependent. No single data format or technology automatically makes a dataset AI-ready.

AI does not introduce a new interoperability layer. Instead, automated AI workflows increase the importance of structural, semantic, and technical interoperability.

Interoperable source data provide a foundation from which task-specific AI training datasets can be constructed reproducibly.

Cloud-native layouts such as Zarr can improve selective and parallel access for suitable large-scale workflows, but they do not automatically improve semantic interoperability or data quality.

Persistent identification, provenance, and versioning complement interoperability by connecting models and derived training data to the exact source data and transformations from which they were produced.

Interoperability supports reproducible and inspectable AI workflows, but it does not guarantee that training data are representative, scientifically appropriate, unbiased, or sufficient for a particular model.