# PAYU issues on Leonardo

**URL:** <https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914>\
**Category:** Technical\
**Tags:** help, payu\
**Created:** [18 November 2024 01:51 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914 "2024-11-18T01:51:34Z")\
**Posts on this page:** 13\
**Page:** 2

<div class="post-metadata">

**Author:** ![Natalia](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/natalia/32/1612_2.png) [@Natalia](https://forum.access-hive.org.au/u/Natalia)\
**Post date:** [18 February 2025 22:09 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/21 "2025-02-18T22:09:35Z")

</div>

I finally made it running.

I’ve added a `--oversubscribe` flag to `mpirun`, so it launches executables even it is not enough resources (not the work case, but made things more clear).

The story is that while slurm allocates the proper resources (`--nodes=8 --ntasks=256 --ntasks-per-node=32`):

```auto
[ntilinin@login02 test]$ squeue -u ntilinin
             JOBID PARTITION NAME USER ST TIME NODES NODELIST(REASON)
          12680859 boost_usr mpi_scri ntilinin R 21:07 8 lrdn[0257,0287,0863,0977,2113,2373,2550,2587]

```

`mpirun` only uses 1 node (tested it with test mpi hello world) as OpenMPI 4.1.4 from ACCESS NRI package was built using `‘—without-slurm'` option. `mpirun` or `srun` would not work properly in this case. The ORTE error was due to ssh communication between nodes not configured/not working.

@Aidan, probably this might contribute to the future plans for slurm adaptation you mentioned - install OpenMPI with slurm so that `srun` command works (`‘—without-slurm'` option disabled).

It runs now on one node only (32 cores) and fails with:

> ==\> NOTE from ocean\_model\_init: reading maskmap information from \> INPUT/ocean\_mask\_table  
> parse\_mask\_table: Number of domain regions masked in ocean model = 24
> 
> FATAL from PE 0: fms\_io(parse\_mask\_table\_2d): mpp\_npes() .NE. layout(1)\*layout(2) - nmask for ocean model

Probably this has something to do with the ocean mask size, the number of MPI processes does not fit 32 cores and requires more.

@angus-g, you gave a right hint, thank you!

@anton, thank you for advices! This was the beginning of my journey last year with ACCESS OM2, I compiled it with system provided OpenMPI and netcdf from the COSIMA repo. That worked somehow, but then the COSIMA repo was no longer supported and migrated to ACCESS NRI and I was advised to use spack, we installed everything with Harshula online from ACCESS NRI package using spack (model \*.exe files, OpenMPI, etc.)  
I’m aware (read somewhere on forum) that spack built installation runs slower than executables compiled with system provided OpenMPI, but at this point I just want to make it working one or the other way.

So now I probably should go back to the beginning and compile the model with system wide OpenMPI and netcdf again. Leornardo supports spack and I can probably install the OpenMPI and netcdf versions that would be compatible with ACCESS OM2 source code, will try to workaround with spack now (helpdesk only works with tickets/e-mails, they are located in another city, might use this option later on too).

Many thanks again for help!

---

<div class="post-metadata">

**Author:** ![anton](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/anton/32/1836_2.png) [@anton](https://forum.access-hive.org.au/u/anton)\
**Post date:** [18 February 2025 23:03 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/22 "2025-02-18T23:03:05Z")

</div>

Glad to hear you are making progress !

---

<div class="post-metadata">

**Author:** ![Aidan](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/aidan/32/42_2.png) [@Aidan](https://forum.access-hive.org.au/u/Aidan)\
**Post date:** [19 February 2025 01:18 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/23 "2025-02-19T01:18:59Z")

</div>

> [@Natalia](#):
>
> So now I probably should go back to the beginning and compile the model with system wide OpenMPI and netcdf again. Leornardo supports spack and I can probably install the OpenMPI and netcdf versions that would be compatible with ACCESS OM2 source code, will try to workaround with spack now

You can use the system OpenMPI by setting mpi as an external non-buildable package, e.g.

> <https://github.com/ACCESS-NRI/spack-config/blob/main/common/gadi/packages.yaml#L11-L329>

Note that this configuration contains a number of versions, but you would not need to specify that many, or your HPC system might provide a similar package file you could refer to or use.

netCDF is straightforward to compile, so I would just build that with spack if you’re happy to do so.

---

<div class="post-metadata">

**Author:** ![Natalia](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/natalia/32/1612_2.png) [@Natalia](https://forum.access-hive.org.au/u/Natalia)\
**Post date:** [28 February 2025 06:04 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/24 "2025-02-28T06:04:07Z")

</div>

It is running on Leonardo. The whole adjustment of OpenMPI and other things was quite painful 😅  
The key issues was related to force OpenMPI (built with spack, as Leonardo doesnt’ have OpenMPI module compiled with intel compilers, only gcc and nvhpc).  
The very specific thing was forcing spack built OpenMPI to use Infiniband to exchange between nodes, by default it was picking the TCP Ethernet. Also the `srun` with heterogeneous jobs.  
Now I have this whole thing running:

```auto
6:53 $ sacct -j 13143246 --format=JobID,JobName,State,Start,End,Elapsed,NCPUS,ReqTRES,AllocTRES
JobID JobName State Start End Elapsed NCPUS ReqTRES AllocTRES 
------------ ---------- ---------- ------------------- ------------------- ---------- ---------- ---------- ---------- 
13143246 wrap RUNNING 2025-02-28T06:11:30 Unknown 00:47:41 560 billing=5+ billing=5+ 
13143246.ba+ batch RUNNING 2025-02-28T06:11:30 Unknown 00:47:41 112 cpu=112,g+ 
13143246.ex+ extern RUNNING 2025-02-28T06:11:30 Unknown 00:47:41 560 billing=5+ 
13143246.0+0 yatm.exe RUNNING 2025-02-28T06:11:44 Unknown 00:47:27 1 cpu=1,gre+ 
13143246.0+1 fms_ACCES+ RUNNING 2025-02-28T06:11:44 Unknown 00:47:27 216 cpu=216,g+ 
13143246.0+2 cice_ausc+ RUNNING 2025-02-28T06:11:44 Unknown 00:47:27 24 cpu=24,gr+

```

It produces the .err and .out files, but MOM keeps producing this output for different ocean 2-d and 3-d fields:

```auto
WARNING from PE 0: diag_util_mod::opening_file: one axis has auxiliary but the corresponding field is NOT found in file ocean-3d-ty_trans-1-monthly-mean-ym%4yr%2mo

WARNING from PE 0: diag_util_mod::opening_file: one axis has auxiliary but the corresponding field is NOT found in file ocean-2d-ty_trans_int_z-1-monthly-mean-ym%4yr%2mo

WARNING from PE 0: diag_util_mod::opening_file: one axis has auxiliary but the corresponding field is NOT found in file ocean-2d-mld-1-monthly-mean-ym%4yr%2mo

```

However it keeps going and not failing.

Maybe this error has been seen already? Would be grateful for any hint on it.

Thanks as usual!

---

<div class="post-metadata">

**Author:** ![anton](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/anton/32/1836_2.png) [@anton](https://forum.access-hive.org.au/u/anton)\
**Post date:** [2 March 2025 22:16 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/25 "2025-03-02T22:16:05Z")

</div>

Good to hear its working !

Has a model run through to completion and then you can restart the model for a second run ?

Are your Payu changes on github ? There may be some lessons for us to integrate into Payu 🙂

> [@Natalia](#):
>
> `WARNING from PE 0: diag_util_mod::opening_file: one axis has auxiliary but the corresponding field is NOT found in file ocean-3d-ty_trans-1-monthly-mean-ym%4yr%2mo`

I think this warning is ok and we also get this warning - it just means that the metadata for a variable is referring to a different variable which is not in that file. This is ok because the variable it is referencing is in a different output file.

---

<div class="post-metadata">

**Author:** ![anton](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/anton/32/1836_2.png) [@anton](https://forum.access-hive.org.au/u/anton)\
**Post date:** [2 March 2025 22:20 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/26 "2025-03-02T22:20:26Z")

</div>



---

<div class="post-metadata">

**Author:** ![Natalia](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/natalia/32/1612_2.png) [@Natalia](https://forum.access-hive.org.au/u/Natalia)\
**Post date:** [3 March 2025 18:46 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/28 "2025-03-03T18:46:34Z")

</div>

Thanks!

Yes, I’ve run 2 years long experiment and then restarted it from the last saved restarts for another 25 years. All works well, can’t believe it 💥

I’ll soon create forks for all repos I modified and push my changes.  
I’ll also sum up all the adjustments in the description file. Will post here all links.

@anton, am I getting right that probably the the easiest way to use ESMValTool on Leonardo to evaluate the model? As all other software is adapted to Gadi and use `access-nri-intake-catalog`?

---

<div class="post-metadata">

**Author:** ![Aidan](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/aidan/32/42_2.png) [@Aidan](https://forum.access-hive.org.au/u/Aidan)\
**Post date:** [3 March 2025 23:23 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/29 "2025-03-03T23:23:06Z")

</div>

> [@Natalia](#):
>
> @anton, am I getting right that probably the the easiest way to use ESMValTool on Leonardo to evaluate the model? As all other software is adapted to Gadi and use `access-nri-intake-catalog`?

If you’re happy to do so I’d suggest making a new topic for this question. Then we can close this one and get some of the model evaluation team involved to assist.

---

<div class="post-metadata">

**Author:** ![anton](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/anton/32/1836_2.png) [@anton](https://forum.access-hive.org.au/u/anton)\
**Post date:** [3 March 2025 23:57 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/30 "2025-03-03T23:57:51Z")

</div>

> [@Natalia](#):
>
> @anton, am I getting right that probably the the easiest way to use ESMValTool on Leonardo to evaluate the model? As all other software is adapted to Gadi and use `access-nri-intake-catalog`?

This is a tricky question somewhat - it depends on what you are trying evaluate and what the goals are … in many ways it also depends on what you colleagues use.

ESMValTool has recipes in it for CMIP style inter-comparison of model results. I think it is most targetted at repeating the same analysis across different modelling centres, and where analysis is targetting the _robustness and confidence in the model results and evaluate the performance of models against observations or against predecessor versions of the same models_ ([Righi et al 2020](https://gmd.copernicus.org/articles/13/1179/2020/))

The `access-nri-intake-catalog` is built on the [intake-esm](https://intake-esm.readthedocs.io/en/stable/) framework. This is a tool for making catalogues - i.e. a catalog for searching and finding data and variables for experiments (and observational) data. The data can be loaded (typically into an [xarray](https://docs.xarray.dev/en/stable/index.html) dataset, but could be something else). The actual analysis then is up to the individual. [cosima-recipes](https://cosima-recipes.readthedocs.io/en/latest/) has examples of using intake-esm to load data into a xarray object. You could use the `builders` from the `access-nri-intake-catalog` to make small intake-esm datastores for your experiments, or you could load you data directly into xarray objects using their file paths and xarray’s `open_mfdataset`

---

<div class="post-metadata">

**Author:** ![Natalia](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/natalia/32/1612_2.png) [@Natalia](https://forum.access-hive.org.au/u/Natalia)\
**Post date:** [4 March 2025 06:08 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/31 "2025-03-04T06:08:35Z")

</div>

Sure! Apologies for overposting.

---

<div class="post-metadata">

**Author:** ![Aidan](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/aidan/32/42_2.png) [@Aidan](https://forum.access-hive.org.au/u/Aidan)\
**Post date:** [4 March 2025 06:10 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/32 "2025-03-04T06:10:37Z")

</div>

Not at all! We’re very excited to have an international collaborator!

It’s just that if we’re pivoting to a new problem a new topic would be a good move.

Tag with #help to make sure the triage team pick it up.

Cheers

Aidan

---

<div class="post-metadata">

**Author:** ![Aidan](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/aidan/32/42_2.png) [@Aidan](https://forum.access-hive.org.au/u/Aidan)\
**Post date:** [23 April 2025 02:35 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/34 "2025-04-23T02:35:31Z")

</div>

I’ll close this now, but please do open a new topic if you need further assistance @Natalia

---

<div class="post-metadata">

**Author:** ![Aidan](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/aidan/32/42_2.png) [@Aidan](https://forum.access-hive.org.au/u/Aidan)\
**Post date:** [23 April 2025 02:35 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/35 "2025-04-23T02:35:36Z")

</div>



[Previous page](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914.md?page=1)
