Experiment Proposal: Coarsening of the fine-scale atmosphere for cross-scale learning

Experiment title :bell:: Coarsening of fine-scale simulations for cross-scale learning

Summary :bell::

We request resources to enable the building of a database of coarse-grained outputs from high-res (LES and km-scale) numerical simulations.

Scientific motivation:

A fundamental unsolved problem in atmospheric sciences is how to relate behaviour across scales, in particular to quantify and understand how behaviour at small scales (e.g. unresolved by a given model) feeds back onto larger scales (resolved by that model). This applies to atmospheric convection, clouds and general turbulence. ML techniques can now help learn this. However, the bottleneck now preventing progress is a lack of usable CRM or LES datasets. To be usable the data must be stored at much higher temporal resolution (30-60 minutes), and with more output variables, than is typically practical. This storage can however be temporary, as once it is post-processed to obtain the required coarse-grained quantities, most of the original data can be deleted.
The concept follows Shen et al. 2022 but with a broader range of boundary conditions and more complete set of coarse-grained fields.

Experiment Name :bell:: Coarsening of fine-scale simulations for cross-scale learning
People :bell:: Steven Sherwood, Nathan Lue and others
Model: WRF and others
Configuration: WRFlux 4.3 idealised with periodic boundaries (Gobel et al. 2021)
Initial conditions: Initial and nudging to local evolving states sampled from a GCM
Run plan: Run three days at 1-km resolution to equilibrate followed by two days at 200-m
Simulation details:
Total KSUs required :bell:: None requested here
Total storage required :bell:: 50 TB
Storage lifetime :bell:: Two months
Long term data plan :bell:: Cull data
Outputs: Coarse-grained fields and tendencies including fine-scale transports of prognostic variables
Restarts: One required (see above)
Related articles: 10.1029/2021ms002631 10.5194/gmd-15-669-2022

Analysis:

Conclusion:

While the above notes only the Sherwood group’s WRF simulations, we have contacted others in Australia who are interested in a similar approach, and a prime target would be km-scale simulations using the ACCESS model. These simulations would need code development for coarse-graining / post-processing, which we hope could be supported at a higher level above this WG by the NRI via a separate proposal. The goal would be to put outputs in a common format. This proposal will provide enough storage to enable initial testing and development.

Thanks for the proposals. The working group co-chairs discussed the proposal and are happy to support the storage utilisation for this period of time. A few notes.

(1) The compute and storage resources are not heavily subscribed at the moment, so for anyone else wondering about the application process, this could change if there is more demand for resources.

(2) It would be great if the project science leads could attend and present on this work at an upcoming ML community WG meeting

(3) It is possible others in the community might like to take an interest in the data. I presume the data would be available to others in the ML WG opportunistically before it is culled.

Thanks Tennessee – of course happy to provide the data, that is our intention.

Hello all,

There will be a breakout session at the upcoming NRI ML meeting in August on the topic of hybrid modelling where we can include some discussion on if and how to progress this project–that is, toward a pipeline for calculating and saving coarse-grained quantities from the fine (km)-scale runs we are already doing, to serve as training data for ML and science projects for linking across scales.
In the mean time–and to allow input from those not attending next month–I would love to hear and for people to start thinking about what capacity/infrastructure and interest there is, and what more would be needed. And maybe have a discussion or Zoom meeting before the ML workshop. Comments very welcome.

Those in the ATM WG who expressed interest were @heidi, @qinggangg , @traupach , @pjs548 , @bethanwhite , @Paul.Gregory , and @Yi_Huang . Others I recall are @VassiliKitsios , @Claire_Vincent , @JulieA and @marty.singh , and my co-conspirators @davidf and @abhnil but I lost my notes on the modelling and ML WGs, so there are probably others and I hope they see this!

Hi Steve,

Thanks for pushing this forward. My only comment is that we should ensure, as best as we can, that we include quantities in the coarse-grained data that are useful for budgets: True averages of things like radiative and surface fluxes, true correlations of u,v and w with relevant variables (temperature, humidity, and horizontal momentum). and ideally perhaps apparent heating/drying tendencies from the subgrid schemes. Not sure how feasible these would be.

HI Marty, yes that is the idea. For WRF we are using a package that does this. For ACCESS it would require development work.