Post processing tool to convert UM files to NetCDF

ACCESS-NRI is seeking community feedback on um2nc, a tool developed by ACCESS-NRI to convert UM (Unified Model) files to NetCDF format.

The tool is currently integrated into post-processing workflows for ACCESS-ESM1.5 and ACCESS-ESM1.6. It is also designed for use in post-processing scripts to convert output from other models, such as ACCESS-rAM3 and ACCESS-AM3, or as a standalone command-line tool.

While it does not feature in the recent ACCESS-AM3-beta release and current ACCESS-rAM3 release, we are planning to include this um2nc tool in the upcoming ACCESS-AM3 full release and the next release of ACCESS-rAM3.

How to get um2nc on Gadi

um2nc is available in the model-processing development environment on Gadi:

module use /g/data/vk83/modules
module load model-processing/2026.09.15

How to use um2nc

For usage documentation, run:

um2nc --help

and more specifically for commands:

um2nc convert --help

There are numerous command-line options to suit different workflows, so we recommend checking this usage help to find what works for your use case.

Different um2nc use-cases

There are two main ways to use um2nc:

  • offline (outside modelling suites/workflows)
  • online (embedded within a modelling suite/workflow)

Use um2nc outside model suites/workflows

This is usually the easiest and most straightforward way to use um2nc.
It requires you to manually use the um2nc convert command on the UM output files generated by the model.

Use um2nc within model suite/workflows

The exact steps for this depend on the specific setup of your model suite or workflow.

The general approach is to modify your Cylc suite/workflow to add or update a post-processing task that runs after the model produces its outputs. In this task, you can run the module use and module load commands to make um2nc available, then run um2nc convert on the UM output files generated by the model. Alternatively, the module use and module load commands can be added as an init-script for the task.

How to provide feedback

We would appreciate feedback on your experience using this um2nc tool on your model output and data, including:

  • Issues or unexpected behaviour
  • Ideas for new features or improvements
  • Suggestions on how we can make it more useful for your workflows
  • Pros and Cons of ACCESS-NRI’s um2nc compared to other similar tools, such as BOM’s netCDF conversion script, old um2netcdf.py script, and UKMO UM file conversion tool.

:backhand_index_pointing_right: Please reply to this topic with any comments, questions, or suggestions.
We’re keen to understand how the ACCESS community is using this tool and how we can make it better!

Hi @atteggiani and @heidi,

Firstly, thank you for this!

Secondly, I use a version of rAM3 in which we have incorporated the Bureau’s um2nc conversion workflow, which also deletes the model’s native fields files after um2nc conversion has been done.

In order to test this version against the Bureau version that I currently use, is there any chance of being given access to an alpha implementation of it in rAM3?

Thanks!
Bethan.

Hi @bethanwhite,

I don’t think there is an implementation of our um2nc in rAM3 yet, but if you share your version of rAM3 I can try and set up a modified version that uses our um2nc instead of the Bureau one :slight_smile:

Cheers
Davide

Thanks Davide!

Our suite is in the fcm repo, which you can get with:
svn co https://code.metoffice.gov.uk/svn/roses-u/b/y/3/9/5/ram3_flagship

If you make a new branch, can you name it something that’s nicely distinct? We’re trying to avoid branch name proliferation & prune old versions out of the repo as we’ve had a few instances of people picking up the wrong copy (older dev branches) to work with because of similar branch names.

Thanks Davide, that would be great.

You might like to try this version instead, which is a direct port of ram3 without all the extra flagship stuff. It should have lower overheads than the flagship suite.

https://code.metoffice.gov.uk/svn/roses-u/b/y/3/9/5/ram3_um2nc

Much better idea :up_arrow:

Hi all,

Sorry for the delay in my reply but I’ve been caught up with other projects.

ACCESS-rAM3 with ACCESS-NRI’s um2nc support

I have set up the nci_access_ram3_um2nc branch of ACCESS-rAM3 with support for NetCDF conversion using ACCESS-NRI’s um2nc.

The branch adds a task to convert all outputs to NetCDF using the “basic” um2nc conversion.
If you would like to test it using different options (e.g., one file per stash → --one-nc-per-stash-variable, different STASH name mapping → --model, etc.) please add the desired options here.

Alternatively, you can also test the conversion “offline”, by loading the um2nc tool:

module use /g/data/vk83/modules
module load model-processing/2026.09.15

and running:

um2nc convert <input> <output> [options]

on the desired output files.

For more information on um2nc options, run:

um2nc convert --help

Provide Feedback

We would really welcome your feedback on this functionality!
Please refer to the introductory post for feedback report.

Thank you! :slight_smile:
Davide

Brilliant, thanks @atteggiani!

I’ve got a bit of remaining Q3 compute that I will use to test this :slightly_smiling_face:

Thanks @atteggiani for putting the um2nc task into the rAM3 suite to allow us to test this from the perspective of ACCESS-rAM3 end users :tada:

This is combined feedback from myself (@bethanwhite) and @mlipson. We’ve tested this out on two different rAM3 configurations.

We find that the out-of-the-box conversion task on the standard Lismore domain run, as well as on a 12km grid-spacing large continental-scale domain, is nice and efficient :white_check_mark:

It is great that the release of this as a suite task contains a namelist / GUI-level switch that is on by default :white_check_mark:

However, to make this implementation useful to rAM3 users it needs specific rAM3 adaptations to be deployed in the main release:

  • The conversion needs to be single netcdf files per STASH variable by default (i.e. the um2nc task script in a rAM3 deployment needs to contain --one-nc-per-stash-variable). Everyone is storage-limited, and using single field files not only makes data processing and analysis much simpler, but also makes longer-term data management easier too.

  • However, the files need to have sensible naming based on their variable names, not STASH code-based naming. Filenames like fld_s30i404_umnsaa_pverb012.nc contain zero useful information and make data management and processing, as well as viewing in e.g. ncview, difficult.

  • The lack of a --model option and associated STASHmaster for rAM3 is a barrier to the above point - for deployment of this um2nc in the main rAM3 release, a STASHmaster needs to be created for the standard output set ACCESS-NRI has chosen for the rAM3 release so that output files are named sensibly, and the rAM3 um2nc task script needs to then have the relevant --model access-ram3 setting as default.

  • By extension of this, there also needs to be a clear path to allow people to add metadata for their own variables, as people will often include new STASH outputs outside of the rAM3 defaults.

  • File naming should also be smarter about aggregation - for example, mean and instantaneous precipitation are both called pr in the file naming, with the mean adding a _1.nc. This obfuscates important information when looking through output directories. More useful filenaming (for this precip ,pr, example) would be e.g. pr_instant, pr_mean and pr_accum.

  • Ideally file naming would lead with the variable name and then contain other info, like the aggregation, as a suffix. In the BoM implementation we have, the aggregation is a prefix, which means that the same variables don’t sit together when you list them in a directory (e.g. instantaneous and mean precip).

  • netcdf files within each cycle should be concatenated to a single cycle, e.g. if running 24 hour cycles with 6-hourly forecast tasks, the .nc files should be concatenated to one file per 24-hour cycle. There’s no reason to have separate 00, 06, 12 and 18 files - this not only makes data management, processing and analysis more difficult, but increases the inode count of the output by a factor of (in this example) 4.

  • Finally, a deployment of this in the main rAM3 release should also include an extra task that deletes the native model fields files output and leaves only the netcdf files. No end user would ever have reason to choose fields files over netcdf, and everyone is storage-constrained. In the current version, the um2nc task doesn’t delete the native model output, so the conversion adds to the total output volume.

  • This removal of the native output files after successful conversion should be done per-forecast cycle rather than one single task at the end of the suite, as this will reduce the data volume overheads required by the suite during run-time.

The BOM um2nc conversion that we have implemented into some of our rAM3 branches contains all these features, so we are able to point you to this if you need examples to work from.

Thanks again to ACCESS-NRI for considering our original request for in-suite netcdf conversion; it is a very important feature that enables more people to run more rAM3 experiments, and to do more with their model outputs.

Bethan & Mat.

To illustrate the file naming, this is what listing the contents of a directory of 1 model day’s output with single file output conversion looks like for vanilla rAM3 output fields.

In addition to the points made above about file naming, there is no need to retain the parent file info (i.e. the umnsaa_pverb000 part) in the filenames - this could go into the metadata instead.

Thanks a lot @bethanwhite and @mlipson for the extensive feedback!
This is very valuable to us for driving the development of um2nc and we’ll work to implement the points you raised.

About file naming, what you would like seems to be something like

<stash_var_name>_<aggregation>.nc

am I correct?

For filename uniqueness I suggest:

<stash_var_name>_<aggregation>_<region>_<resolution>_<model>_YYYYMMDDHHMM-YYYYMMDDHHMM.nc

This is based on the BoM um2nc (although they include a version number, which I don’t think rAM3 would need):

netswsfc-NA_0p11_GAL9-v1-200010040030-200010042330.nc

An additional functionality of BoMs implementation I found useful was the ability to define a location on scratch or gdata outside the suite for processed netcdfs. This was defined with: NCI_STORAGE_TOP_DIR in rose-suite.conf. Separating output from the rest of the suite data means that we can easily delete the suite directory without worrying we are removing outputs, important as we are highly storage constrained.

Finally, the BoM implementation had separate tasks for each “stream” or file output, with resources (cpu/memory) being configurable for each stream. This is useful for large domains and high resolution 3D outputs (e.g. the centre’s flagship) which require large resources to process. I’m not sure if this is necessary with the NRI implementation, I haven’t yet tested it with a large domain or I/O heavy outputs.

Thank you for the reply @mlipson.

I agree it would be better to make the output files unique.
The format you listed here sounds good to me.

I think separating the output from the rest of the suite is a good idea.
Instead of using the NCI_STORAGE_TOP_DIR config parameter which already exists and might be needed for other purposes, I would add a separate config parameter specifically for the NetCDF output directory. For example: NETCDF_OUTPUT_DIR.

Not sure about this one. For simplicity, I would try to keep only one task for all different streams. The idea for the final implementation is to allow um2nc to get all the information about the outputs (streams, output file names, etc.) by reading the UM namelist itself, so we can be more consistent and precise about the files we need to process (also if the user changes the output settings). A similar functionality (not using um2nc yet though) is already being tested for ACCESS-AM3.

In any case, the current um2nc is not optimised to be memory-efficient and I think improvements could made in that sense. I think it would be good to carry out some tests using large domains with heavy output files, to assess how much memory optimisation we need or if we can be satisfied with the current functionality.

Cheers

Thanks @atteggiani!

On the last point re memory optimisation on large domains with heavy output files, I’m currently running a 1km entire north-Aus tropical channel domain for a TC data vis project with ACCESS-NRI. This is with our Flagship suite, but happy to share conf files and ancils with you if you would like to use the same domain for testing.

Thanks @bethanwhite, that would be useful! :slight_smile:

Hi all,

Thanks for the feedback and suggestions for um2nc!

  • I agree that variable names are not very useful at the moment. We’d hoped to have a better list of names for the recent release of ESM1.6 but the mapping of STASH codes to useful names was not easy. Taking another look at this in general, not just for rAM3, is definitely on the to do list!
  • For filenames I would like to propose we follow the naming convention that is being used for OM3 and ESM1.6:
    • <model>.<component>.<dimension>.<field>.<frequency>.<time_cell_method>.<datestamp>.nc
    • For rAM3 this would look something like access-ram3.um13p5.2d.netswsfc.1hr.mean.20001004.nc
    • We might need to tweak some details for rAM3 - we might need starttime & endtime instead of datestamp, and it might be worth adding a region portion for non-global models.
    • It’s not perfect, there are some details I would change if we were starting from scratch, but being consistent with the other ACCESS model output is worthwhile in my opinion

To be a bit stronger than @joshuatorrance . When he says he’d like to propose the following naming convention, it means ACCESS-NRI will use this naming convention and the output data specification shared by @joshuatorrance in future releases of ACCESS model configurations we support. Some modifications (as highlighted by Joshua) might be needed for some models and will be worked on. When the post-processing will be modified to apply the new output specification to each model is not yet completely determined, but all models will move to a similar specification for output files eventually.

Consitency in output files will facilitate research using several ACCESS models.

We welcome any feedback on the specification and involvement from the community on the portions that are still in discussion (such as the variable names to replace UM field names and modifications to the filenames for regional models for example) as well as feedback on the post-processing tool itself.

Hi Josh, that sounds good to align with other ACCESS models.

For rAM3 an important inclusion would be resolution, as we typically run 2-3 nested simulations at different resolutions which would otherwise have the same filename.

This can be taken from the rAM3 suite, as the resolution names are defined values.

The same goes for region and model config (e.g. if we are doing two seperate experiments at the same resolution), but I can understand that makes the naming quite long, and that could be managed through having seperate folders, so I could be convinced to drop that request (although I still find totally unique filenames useful).

Hi @joshuatorrance,

it might be worth adding a region portion for non-global models.

I agree with @mlipson - for rAM3 naming it would be preferable to pull in information about the region, resolution and configuration.

Taking the rAM3 Lismore case default example, the rose-suite.conf contains the following (user-configurable) information:

rg01_name="Lismore"
rg01_nreslns=2
rg01_rs01_name="d1100"
rg01_rs02_name="d0198"
rg01_rs01_m01_name="GAL9"
rg01_rs02_m01_name="RAL3P3"

This could be used to create the following filename contributions by extracting and concatenatingrg0X_name.rg0X_rs0Y_name.rg0X_rs0Y_m0Z_name

In the Lismore example, this would mean that files produced by the Lismore coarse domain running GAL9 science (and named to reflect this) would contain in their filename:.Lismore.d1100.GAL9.

And files produced by the nested higher resolution Lismore domain running RAL3.3 science (named to reflect this) would contain:.Lismore.d0198.RAL3P3.

This ensures unique filenaming for data produced by each nested grid, as well as data provenance in the event that someone deletes their suite from the repository for whatever reason, but keeps the data.

we might need starttime & endtime instead of datestamp

I’d be happy with just the datestamp of the first time in the output file - I don’t see a need to include end time.

Ideally files would be aggregated up to daily.

I’m not certain how the configuration information will work for rAM3 with CABLE. Is that informative to have GAL9 in the filename when it isn’t a pure GAL9 configuration?

Same thing if anyone runs a personal configuration but leaves GAL9 and RAL3P3 in the configuration file. Would we end up with a bunch of files that aren’t quite what they are saying they are?