# PAYU issues on Leonardo

**URL:** <https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914>\
**Category:** Technical\
**Tags:** help, payu\
**Created:** [18 November 2024 01:51 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914 "2024-11-18T01:51:34Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Natalia](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/natalia/32/1612_2.png) [@Natalia](https://forum.access-hive.org.au/u/Natalia)\
**Post date:** [18 November 2024 01:51 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/1 "2024-11-18T01:51:34Z")

</div>

Hello to everyone,

From what I found here, probably this question should be addressed to @Aidan @john_reilly @angus-g @dale.roberts @harshula

I have some progress with payu and porting ACCESS OM2 to Leonardo supercomputer.

Right now at the stage of **payu run**.

**Long story short:**

1. \*.exe files are compiled with ACCESS NRI local modules (compiled with spack)
2. RYF JRA-55 files calculated locally on Leonardo
3. Initial conditions (transferred to Leonardo) and forcing fields specified in config.yaml and atmosphere/forcing.json

payu setup produced manifests, can be found in my local repo: [GitHub - VanuatuN/1deg\_jra55\_ryf: 1 degree ACCESS-OM2 experiment with JRA55 RYF atmospheric forcing.](https://github.com/VanuatuN/1deg_jra55_ryf)

**The questions are:**

- how to force payu to use my local modules that were compiled at the first stage?  
I know it is going to be slower than with system modules, but I want to make it just  
working first
- where exactly in payu/\*.py files I should modify the rest of the slurm specific flags for Leonardo?

**Example of batch script:**  
#!/bin/bash  
#SBATCH --job-name=benchmark\_test  
#SBATCH --output=benchmark\_test.out  
#SBATCH --error=benchmark\_test.err  
#SBATCH --nodes=1  
#SBATCH --cpus-per-task=32  
#SBATCH -A ICT24\_MHPC  
#SBATCH --time=00:30:00  
#SBATCH --partition=boost\_usr\_prod

Thank you!

The current output from payu run:

```auto
02:48 $ payu run 
payu: warning: Job request includes 47 unused CPUs.
payu: warning: CPU request increased from 241 to 288
sbatch -A ICT24_MHPC --time=10800 --ntasks=288 --wrap="/leonardo/prod/spack/5.2/install/0.21/linux-rhel8-icelake/
gcc-8.5.0/anaconda3-2023.09-0-zcre7pfofz45c3btxpdk5zvcicdq5evx/bin/
python /leonardo/home/userexternal/ntilinin/.local/bin/payu-run" --export="PAYU_PATH=/leonardo/home/userexternal/ntilinin/.local/bin,MODULESHOME
=/leonardo/prod/spack/03/install/0.19/linux-rhel8-icelake/gcc-8.5.0/
environment-modules-5.2.0-rz47odw4phlhzhhbz7b65nv5s5othgmi,MODULES_CMD=/leonardo/prod/spack/03/install/
0.19/linux-rhel8-icelake/gcc-8.5.0/environment-modules-5.2.0-rz47odw4phlhzhhbz7b65nv5s5othgmi/libexec/modulecmd.tcl,MODULEPATH=
/leonardo/prod/spack/03/install/0.19/linux-rhel8-icelake/gcc-8.5.0/environment-modules-5.2.0-rz47odw4phlhzhhbz7b65nv5s5othgmi/modulefiles:/leonardo/prod/opt/modulefiles/
profiles:/leonardo/prod/opt/modulefiles/base/archive:/leonardo/prod/opt/modulefiles/
base/dependencies:/leonardo/prod/opt/modulefiles/base/data:/leonardo/prod/opt/
modulefiles/base/environment:/leonardo/prod/opt/modulefiles/base/libraries:/leonardo/
prod/opt/modulefiles/base/tools:/leonardo/prod/opt/modulefiles/base/compilers:/leonardo/prod/opt/modulefiles/base/applications"
sbatch: error: no partition specified, using default partition lrd_all_serial
sbatch: error: no gres:tmpfs specified, using default: gres:tmpfs:10g
sbatch: error: Batch job submission failed: More processors requested than permitted
Traceback (most recent call last):
  File "/leonardo/home/userexternal/ntilinin/.local/bin/payu", line 10, in <module>
    sys.exit(parse())
             ^^^^^^^
  File "/leonardo/home/userexternal/ntilinin/.local/lib/python3.11/site-packages/payu/cli.py", line 42, in parse
    run_cmd(**args)
  File "/leonardo/home/userexternal/ntilinin/.local/lib/python3.11/site-packages/payu/subcommands/run_cmd.py", line 108, in runcmd
    cli.submit_job('payu-run', pbs_config, pbs_vars)
  File "/leonardo/home/userexternal/ntilinin/.local/lib/python3.11/site-packages/payu/cli.py", line 156, in submit_job
    subprocess.check_call(shlex.split(cmd))
  File "/leonardo/prod/spack/5.2/install/0.21/linux-rhel8-icelake/gcc-8.5.0/anaconda3-2023.09-0-zcre7pfofz45c3btxpdk5zvcicdq5evx/lib/python3.11/subprocess.py", line 413, in check_call
    raise CalledProcessError(retcode, cmd)
subprocess.CalledProcessError: Command '['sbatch', '-A', 'ICT24_MHPC', '--time=10800', '--ntasks=288',
 '--wrap=/leonardo/prod/spack/5.2/install/0.21/linux-rhel8-icelake/gcc-8.5.0/
anaconda3-2023.09-0-zcre7pfofz45c3btxpdk5zvcicdq5evx/bin/python 
/leonardo/home/userexternal/ntilinin/.local/bin/payu-run', '--export=PAYU_PATH=/leonardo/home/userexternal/ntilinin/.local/
bin,MODULESHOME=/leonardo/prod/spack/03/install/0.19/linux-rhel8-icelake/
gcc-8.5.0/environment-modules-5.2.0-rz47odw4phlhzhhbz7b65nv5s5othgmi,MODULES_CMD=
/leonardo/prod/spack/03/install/0.19/linux-rhel8-icelake/gcc-8.5.0/
environment-modules-5.2.0-rz47odw4phlhzhhbz7b65nv5s5othgmi/libexec/modulecmd.tcl,
MODULEPATH=/leonardo/prod/spack/03/install/0.19/linux-rhel8-icelake/gcc-8.5.0/environment-modules-5.2.0-rz47odw4phlhzhhbz7b65nv5s5othgmi/modulefiles:
/leonardo/prod/opt/modulefiles/profiles:/leonardo/prod/opt/modulefiles/
base/archive:/leonardo/prod/opt/modulefiles/base/dependencies:/leonardo/prod/opt/modulefiles/base/data:/leonardo/prod/opt/modulefiles/base/environment:
/leonardo/prod/opt/modulefiles/base/libraries:/leonardo/prod/opt/modulefiles/base/tools:/leonardo/prod/opt/modulefiles/base/compilers:/leonardo/prod/opt/modulefiles/base/applications']' returned non-zero exit status 1.

```

---

<div class="post-metadata">

**Author:** ![john\_reilly](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/john_reilly/32/237_2.png) [@john\_reilly](https://forum.access-hive.org.au/u/john_reilly)\
**Post date:** [18 November 2024 02:17 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/2 "2024-11-18T02:17:42Z")

</div>

Hi Natalia,

I’m no expert on this stuff and just picked up where @angus-g and @ChrisC28 got to with our slurm-based hpc. But hopefully this helps;

Not sure about the first question sorry, but for the slurm specific flags, we put them at the start of the `config.yaml` file in the run directory. If you can find the `slurm.py` file in `payu/schedulers/` that should make more sense about how payu reads these flags in.

Here’s an example of one of our config.yaml files:

```auto
scheduler: slurm
project: pawsey0410
walltime: 02:20:00
jobname: eac_sthpac-forced_v3
ncpus: 1804
nnodes: 15
runspersub: 1

shortpath: /scratch/pawsey0410
model: mom6
input:
    - /scratch/pawsey0410/jreilly/mom6-inputs/eac_sthpac-forced_v2/
    - /scratch/pawsey0410/jreilly/jra_padded/2016/
    - /scratch/pawsey0410/jreilly/mom6/archive/eac_sthpac-forced_v3/restart305
# - /g/data/ua8/JRA55-do/RYF/v1-3/
# - /g/data/ik11/inputs/JRA-55/RYF/v1-3/
# release exe
exe: /software/projects/pawsey0410/cc7576/mom6-cmake/coupler/MOM6-SIS2
  #exe: /software/projects/pawsey0410/jreilly/mom6-cmake/coupler/MOM6-SIS2
  # /software/projects/pawsey0410/cc7576/mom6-cmake/coupler/MOM6-SIS2

stacksize: unlimited

collate: false
runlog: false

mpi:
  runcmd: srun

```

---

<div class="post-metadata">

**Author:** ![anton](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/anton/32/1836_2.png) [@anton](https://forum.access-hive.org.au/u/anton)\
**Post date:** [18 November 2024 02:53 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/3 "2024-11-18T02:53:40Z")

</div>

Hi Natalia

This looks like where it failed:

> [@Natalia](#):
>
> ```auto
> sbatch: error: Batch job submission failed: More processors requested than permitted
> 
> ```

It looks like the number of nodes is hardcoded to 1, and then payu is requesting 288 cores. I don’t know how many cores per node Leonardo hardware has, for us its 48, so setting the number of nodes to 6 would be correct for us. I would try setting ncpus and nnodes in the config.yaml per Johns code snippet…

There’s these two lines in the payu output:

```auto
payu: warning: Job request includes 47 unused CPUs.
payu: warning: CPU request increased from 241 to 288

```

I think there might be 32 cores per node for you, so I would try:

```auto
ncpus: 256
nnodes: 8
npernode: 32

```

For our normal gadi scheduler, we don’t specify the number of nodes, so its possible Payu hasn’t been tested very well in these cases.

Re: modules

You can set it similar to this config:

> <https://github.com/ACCESS-NRI/access-om3-configs/blob/01cbd6fceb970ab79ee82490df3d4bf8d41447e8/config.yaml#L47-L51>

When you have module: use: and load: lines in the `config.yaml`, you should be able to access the binaries without a path

for example, the `exe:` entry just could become `yatm.exe`

I would test these in a command prompt first by doing a module use and module load, and seeing if the executables are available as commands.

---

<div class="post-metadata">

**Author:** ![anton](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/anton/32/1836_2.png) [@anton](https://forum.access-hive.org.au/u/anton)\
**Post date:** [18 November 2024 04:23 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/4 "2024-11-18T04:23:31Z")

</div>



---

<div class="post-metadata">

**Author:** ![Natalia](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/natalia/32/1612_2.png) [@Natalia](https://forum.access-hive.org.au/u/Natalia)\
**Post date:** [19 November 2024 00:01 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/5 "2024-11-19T00:01:40Z")

</div>

Hi @john_reilly, very much appreciated!  
Didn’t know that slurm flags can be specified exactly in config.yaml

Will try to implement it.

---

<div class="post-metadata">

**Author:** ![Natalia](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/natalia/32/1612_2.png) [@Natalia](https://forum.access-hive.org.au/u/Natalia)\
**Post date:** [19 November 2024 00:03 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/6 "2024-11-19T00:03:45Z")

</div>

Hi @anton!

Thank you, all clear for the moment!  
My time zone forces for a delay in reply.

Very useful information, will do that and let you know soon.

Fingers crossed.

---

<div class="post-metadata">

**Author:** ![anton](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/anton/32/1836_2.png) [@anton](https://forum.access-hive.org.au/u/anton)\
**Post date:** [19 November 2024 00:26 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/7 "2024-11-19T00:26:32Z")

</div>

Happy to help. If the instructions don’t make sense , I can make a Pull Request into your fork of the om2 configurations

---

<div class="post-metadata">

**Author:** ![Natalia](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/natalia/32/1612_2.png) [@Natalia](https://forum.access-hive.org.au/u/Natalia)\
**Post date:** [19 November 2024 20:34 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/8 "2024-11-19T20:34:43Z")

</div>

It worked for the job submission! But failed to pick the modules.  
Leonardo has 32 cores per node, right.  
I modified slurm.py and config.yaml files

**payu run gives:**

```auto
payu run 
sbatch -A ICT24_MHPC --time=00:30:00 --ntasks=256 --partition=boost_usr_prod 
--wrap="/leonardo/prod/spack/5.2/install/0.21/linux-rhel8-icelake/gcc-8.5.0/
anaconda3-2023.09-0-zcre7pfofz45c3btxpdk5zvcicdq5evx/bin/
python /leonardo/home/userexternal/ntilinin/.local/bin/payu-run" --export="PAYU_PATH=/leonardo/home/userexternal/ntilinin/.local/bin,
MODULESHOME=/leonardo/prod/spack/03/install/0.19/linux-rhel8-icelake/
gcc-8.5.0/environment-modules-5.2.0-rz47odw4phlhzhhbz7b65nv5s5othgmi,MODULES_CMD=/leonardo/prod/spack/03/
install/0.19/linux-rhel8-icelake/gcc-8.5.0/environment-modules-5.2.0-rz47odw4phlhzhhbz7b65nv5s5othgmi/libexec/modulecmd.tcl,MODULEPATH=
/leonardo_scratch/large/userexternal/ntilinin/ACCESS-NRI/release/
modules/linux-rhel8-x86_64:/leonardo/prod/spack/03/install/0.19/
linux-rhel8-icelake/gcc-8.5.0/environment-modules-5.2.0-rz47odw4phlhzhhbz7b65nv5s5othgmi/modulefiles:/leonardo/prod/opt/modulefiles/
profiles:/leonardo/prod/opt/modulefiles/base/archive:/leonardo/prod/opt/modulefiles/
base/dependencies:/leonardo/prod/opt/modulefiles/base/data:/leonardo/prod/opt/
modulefiles/base/environment:/leonardo/prod/opt/modulefiles/base/libraries:/leonardo/
prod/opt/modulefiles/base/tools:/leonardo/prod/opt/modulefiles/base/compilers:
/leonardo/prod/opt/modulefiles/base/applications"

```

**But still an output from payu run is:**

```auto
laboratory path: ./ntilinin/access-om2
binary path: ./ntilinin/access-om2/bin
input path: ./ntilinin/access-om2/input
work path: ./ntilinin/access-om2/work
archive path: ./ntilinin/access-om2/archive
nruns: 1 nruns_per_submit: 1 subrun: 1
Loading input manifest: manifests/input.yaml
Loading restart manifest: manifests/restart.yaml
Loading exe manifest: manifests/exe.yaml
Setting up atmosphere
Setting up ocean
Setting up ice
Setting up access-om2
Checking exe and input manifests
Updating full hashes for 3 files in manifests/exe.yaml
Creating restart manifest
Writing manifests/restart.yaml
Writing manifests/exe.yaml
payu: Found modules in /leonardo/prod/spack/03/install/0.19/linux-rhel8-icelake/gcc-8.5.0/environment-modules-5.2.0-rz47odw4phlhzhhbz7b65nv5s5othgmi
Traceback (most recent call last):
  File "/leonardo/home/userexternal/ntilinin/.local/bin/payu-run", line 10, in <module>
    sys.exit(runscript())
             ^^^^^^^^^^^
  File "/leonardo/home/userexternal/ntilinin/.local/lib/python3.11/site-packages/payu/subcommands/run_cmd.py", line 132, in runscript
    expt.run()
  File "/leonardo/home/userexternal/ntilinin/.local/lib/python3.11/site-packages/payu/experiment.py", line 517, in run
    mpi_module = envmod.lib_update(
                 ^^^^^^^^^^^^^^^^^^
  File "/leonardo/home/userexternal/ntilinin/.local/lib/python3.11/site-packages/payu/envmod.py", line 114, in lib_update
    mod_name, mod_version = fsops.splitpath(lib_path)[2:4]
    ^^^^^^^^^^^^^^^^^^^^^
ValueError: not enough values to unpack (expected 2, got 0)

```

**modified `slurm.py` file:**

> <https://github.com/VanuatuN/payu/blob/7372eb99e1a1b04e405203e5c053b5a12c87d357/payu/schedulers/slurm.py#L40-L50>

**modified (with nodes, etc. specified)`config.yaml` file:**

> <https://github.com/VanuatuN/1deg_jra55_ryf/blob/34d8b24129c143043e3613a61ce25cfae5b00522/config.yaml#L10-L20>

Will try to work around with modules in the coming days. But  
would very grateful for any hints where to move.

Thank you!!!

---

<div class="post-metadata">

**Author:** ![anton](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/anton/32/1836_2.png) [@anton](https://forum.access-hive.org.au/u/anton)\
**Post date:** [19 November 2024 21:53 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/9 "2024-11-19T21:53:27Z")

</div>

It looks like payu is trying to check that the mpi version which is linked to by the model executable its the version loaded. But for whatever reason the formatting or check is failing.

I would try adding these lines to your config.yaml and set them to the modules which are used by your exectuables. (You might be able to confirm the path to the mpi version using `ldd` )

```auto
mpi:
    modulepath:
    module:

```

See this section in the docs:

[https://payu.readthedocs.io/en/stable/config.html#miscellaneous](https://payu.readthedocs.io/en/stable/config.html#miscellaneous)

Pinging @Aidan as he has more experience than I with this!

---

<div class="post-metadata">

**Author:** ![Aidan](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/aidan/32/42_2.png) [@Aidan](https://forum.access-hive.org.au/u/Aidan)\
**Post date:** [19 November 2024 22:44 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/10 "2024-11-19T22:44:48Z")

</div>

> [@anton](#):
>
> It looks like payu is trying to check that the mpi version which is linked to by the model executable its the version loaded. But for whatever reason the formatting or check is failing.

Yes this was always quite NCI specific, and with the `spack` built executables is no longer strictly necessary.

Can you try updating your version of `payu`, as there is now logic that isolates this check to NCI systems by matching the library path:

> <https://github.com/payu-org/payu/blob/master/payu/envmod.py#L118>

If you have made local changes you can fetch the latest `payu` and `git rebase` your changes on top of them.

---

<div class="post-metadata">

**Author:** ![Natalia](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/natalia/32/1612_2.png) [@Natalia](https://forum.access-hive.org.au/u/Natalia)\
**Post date:** [24 November 2024 18:30 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/11 "2024-11-24T18:30:24Z")

</div>

**Update:**

Before I was using **‘pawsey’** branch from the payu repo (found somewhere is the issues or here on forum). Now switched to **‘master’** branch and made few corrections.

The job goes to submission, which is goods news.

My opempi module does not recognise --chdir, so I commented out this string and kept -wdir:

> <https://github.com/VanuatuN/payu/blob/aa1f7edc01ff94dee23575d61e5837459c2f61b8/payu/experiment.py#L582-L586>

What I’m not sure about is that payu uses the correct version of as the cmd still looks as:

```auto
 ~/access-om2/control/1deg_jra55_ryf [master ↑·13|…28] 
15:57 $ payu run 
/leonardo/home/userexternal/ntilinin/.local/lib/python3.11/site-packages/payu/fsops.py:77: UserWarning: Duplicate key found in config.yaml: key 'jobname' with value 'access_om2_ryf'. This overwrites the original value: '1deg_jra55_ryf'
/leonardo/home/userexternal/ntilinin/.local/lib/python3.11/site-packages/payu/fsops.py:77: UserWarning: Duplicate key found in config.yaml: key 'queue' with value 'boost_usr_prod'. This overwrites the original value: 'boost_usr_prod'
sbatch -A ICT24_MHPC --time=00:30:00 --ntasks=256 --partition=boost_usr_prod 
--wrap="/leonardo/prod/spack/5.2/install/0.21/linux-rhel8-icelake/
gcc-8.5.0/anaconda3-2023.09-0-zcre7pfofz45c3btxpdk5zvcicdq5evx/
bin/python /leonardo/home/userexternal/ntilinin/.local/bin/payu-run" 
--export="PAYU_PATH=/leonardo/home/userexternal/ntilinin/.local/bin,
MODULESHOME=/leonardo/prod/spack/03/install/0.19/linux-rhel8-icelake/gcc-8.5.0/environment-modules-5.2.0-rz47odw4phlhzhhbz7b65nv5s5othgmi,
MODULES_CMD=/leonardo/prod/spack/03/install/0.19/linux-rhel8-icelake/gcc-8.5.0/environment-modules-5.2.0-rz47odw4phlhzhhbz7b65nv5s5othgmi/libexec/modulecmd.tcl,
MODULEPATH=/leonardo_scratch/large/userexternal/ntilinin/ACCESS-NRI/release/modules/linux-rhel8-x86_64:
/leonardo/prod/spack/03/install/0.19/linux-rhel8-icelake/gcc-8.5.0/environment-modules-5.2.0-rz47odw4phlhzhhbz7b65nv5s5othgmi/modulefiles:
/leonardo/prod/opt/
modulefiles/profiles:
/leonardo/prod/opt/modulefiles/base/archive:
/leonardo/prod/opt/modulefiles/base/dependencies:
/leonardo/prod/opt/modulefiles/base/data:
/leonardo/prod/opt/modulefiles/base/environment:
/leonardo/prod/opt/modulefiles/base/libraries:
/leonardo/prod/opt/modulefiles/base/tools:
/leonardo/prod/opt/modulefiles/base/compilers:
/leonardo/prod/opt/modulefiles/base/applications"
Submitted batch job 9594255

```

It still sees MODULES\_CMD and use systemwide modulecmd.tcl as well as systemwide MODULESHOME, not sure it affects something, but still.

However MODULEPATH is updated to the proper location.

Slurm out looks reasonable, hope it picks not default systemwide openmpi, but first in the list:

> <https://github.com/VanuatuN/1deg_jra55_ryf/blob/a4b5e55151eb1fceec779caa19e25230453be8ee/slurm-9597369.out#L20-L49>

**Another issue** : I’m doing something wrong with resources allocation,  
don’t know how payu distributes submodels across nodes, here is an error that I’m getting now:

> <https://github.com/VanuatuN/1deg_jra55_ryf/blob/a4b5e55151eb1fceec779caa19e25230453be8ee/access-om2.err>

And the config.yaml:

> <https://github.com/VanuatuN/1deg_jra55_ryf/blob/a4b5e55151eb1fceec779caa19e25230453be8ee/config.yaml#L51-L60>

I’ll try to work around, but would be grateful for any advice as usual 🙂

Many thanks!!!

---

<div class="post-metadata">

**Author:** ![anton](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/anton/32/1836_2.png) [@anton](https://forum.access-hive.org.au/u/anton)\
**Post date:** [24 November 2024 21:53 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/12 "2024-11-24T21:53:30Z")

</div>

Hi Natalia

I would try commenting out these lines:

> <https://github.com/payu-org/payu/blob/27aac372dba2752e1d58ab1d66db799e7891e103/payu/experiment.py#L606-L608>

For reasons that are not clear to me, for some reason payu is requesting 16 tasks per available “socket”, when its probably only possible to have one. Maybe this is a gadi specific detail for some specific case. You might be able to remove the `-map-by` argument entirely. I think it will take some experimentation.

---

<div class="post-metadata">

**Author:** ![Natalia](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/natalia/32/1612_2.png) [@Natalia](https://forum.access-hive.org.au/u/Natalia)\
**Post date:** [24 November 2024 22:09 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/13 "2024-11-24T22:09:17Z")

</div>

Hi Anton,

Will try, thank you.  
It could be that each gadi’s node has 3 sockets each with 16 cores, this makes sense - 16\*3=48 cores.

---

<div class="post-metadata">

**Author:** ![Natalia](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/natalia/32/1612_2.png) [@Natalia](https://forum.access-hive.org.au/u/Natalia)\
**Post date:** [14 February 2025 15:12 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/14 "2025-02-14T15:12:56Z")

</div>

Hi @anton

I have an update and a question again.

So far I’ve been struggling during the last month to make payu run executables built from old COSIMA repo with spack built modules (from ACCESS NRI). I rewrote parts of payu that checks modules and adds libraries, but couldn’t make it use proper mpi. Gave up that.

Now I cloned the latest version of payu and trying to run spack built executables with spack built model components from cloned config repo [GitHub - ACCESS-NRI/access-om2-configs at release-1deg\_jra55\_ryf](https://github.com/ACCESS-NRI/access-om2-configs/tree/release-1deg_jra55_ryf). The problem I’m facing now is that one:

mpirun was unable to launch the specified application as it could not access  
or execute an executable:

#--------------------------------------------------------------------------  
mpirun was unable to launch the specified application as it could not access  
or execute an executable:

Executable: ./ntilinin/access-om2/work/1deg\_jra55\_ryf-expt-3216c7cb/atmosphere/yatm.exe  
Node: lrdn3421

while attempting to start process rank 0.  
#--------------------------------------------------------------------------

“which mpirun” points to the proper module prebuilt with spack.  
The symlink points to the proper location of yamt.exe

I do have a feeling that it again uses systemwide mpirun (just a guess).

I found the same issue raised by @Aidan here [ACCESS-OM2 Restart Reproducibility: Bitwise Reproducibility Testing](https://forum.access-hive.org.au/t/access-om2-restart-reproducibility-bitwise-reproducibility-testing/1960)

Maybe there is something specific you and @Aidan can advise on that issues?

Many thanks as usual!

Also an update on nodes/sockets/cores:

It should be 2 sockets per node on Gadi

And it is 1 socket with 32 cores on each Leonardo node, so no need to divide by 2. Payu was configured to divide 32/2 in case of even number ‘npernode’ from config.yaml

---

<div class="post-metadata">

**Author:** ![Natalia](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/natalia/32/1612_2.png) [@Natalia](https://forum.access-hive.org.au/u/Natalia)\
**Post date:** [15 February 2025 10:49 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/15 "2025-02-15T10:49:26Z")

</div>

This is solved now.

The problem was these parts of the code:

> <https://github.com/payu-org/payu/blob/d0fdbff1ca9b52c687d31dbb7773752b0b5223f9/payu/experiment.py#L582-L586>

And here:

> <https://github.com/payu-org/payu/blob/d0fdbff1ca9b52c687d31dbb7773752b0b5223f9/payu/experiment.py#L627-L629>

When using **-wdir** argument with slurm and mpi (probably Leonardo version of slurm, but more likely in general) the line resulting from run\_cmd looks as:

`mpirun --mca io ompio --mca io_ompio_num_aggregators 1 -wdir ./ntilinin/access-om2/work/1deg_jra55_ryf-expt-a3822e12/atmosphere -n 1 ./ntilinin/access-om2/work/1deg_jra55_ryf-expt-a3822e12/atmosphere/yatm.exe`

and after working directory has been changed to:

`/ntilinin/access-om2/work/1deg_jra55_ryf-expt-a3822e12/atmosphere`

it looks for the path to \*.exe files starting from there. So the relative path won’t work.  
I’ve changed the line 629 in experiment.py to

> model\_prog.append(os.path.abspath(os.path.join(model.work\_path, model.exec\_name)))

And mpirun picked to file.

Now I’m facing that output in the access-om2.err:

```auto
--------------------------------------------------------------------------
ORTE has lost communication with a remote daemon.

  HNP daemon : [[55189,0],0] on node lrdn0001
  Remote daemon: [[55189,0],1] on node lrdn0009

This is usually due to either a failure of the TCP network
connection to the node, or possibly an internal failure of
the daemon itself. We cannot recover from this failure, and
therefore will terminate the job.
--------------------------------------------------------------------------
forrtl: error (78): process killed (SIGTERM)
Image PC Routine Line Source             
fms_ACCESS-OM.x 0000000001D5078B Unknown Unknown Unknown
libpthread-2.28.s 000014F4C28ECCF0 Unknown Unknown Unknown
libopen-rte.so.40 000014F4BE89CD30 orte_dt_init Unknown Unknown
libopen-rte.so.40 000014F4BE8E4BB9 orte_ess_base_std Unknown Unknown
libopen-rte.so.40 000014F4BE8E8AA2 Unknown Unknown Unknown
libopen-rte.so.40 000014F4BE95CF0B orte_init Unknown Unknown
libmpi.so.40.30.4 000014F4C3101CE1 ompi_mpi_init Unknown Unknown
libmpi.so.40.30.4 000014F4C2F2F40D MPI_Init Unknown Unknown
libmpi_mpifh.so.4 000014F4C3440787 PMPI_Init_f08 Unknown Unknown
fms_ACCESS-OM.x 0000000001D23B41 coupler_mod_mp_co 82 coupler.F90
fms_ACCESS-OM.x 000000000041F8A0 MAIN__ 186 ocean_solo.F90
fms_ACCESS-OM.x 00000000004111E2 Unknown Unknown Unknown
libc-2.28.so 000014F4C254ED85 __libc_start_main Unknown Unknown
fms_ACCESS-OM.x 00000000004110EE Unknown Unknown Unknown
forrtl: error (78): process killed (SIGTERM)

```

There is still a manual memory allocation in run\_cmd.py in the lines 36-39, that I changed to Leonardo specs:

> ```
> # TODO: Create drivers for servers
> platform = pbs_config.get('platform', {})
> max_cpus_per_node = platform.get('nodesize', 32)
> max_ram_per_node = platform.get('nodemem', 514)
> 
> ```

Probably it gets overwritten from config.yaml, but still.  
Don’t know here to dig now for ORTE error.

Would be grateful for any guess as usual!

---

<div class="post-metadata">

**Author:** ![anton](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/anton/32/1836_2.png) [@anton](https://forum.access-hive.org.au/u/anton)\
**Post date:** [17 February 2025 00:33 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/16 "2025-02-17T00:33:42Z")

</div>

Glad to hear you are making some progress!

Not sure I have much to offer.

> [@Natalia](#):
>
> Don’t know here to dig now for ORTE error.

Yes - this is strange. Can you confirm which code versions you are using ?

It looks like its failing at MPI\_Init for MOM

> <https://github.com/ACCESS-NRI/libaccessom2/blob/f9f2d67a44554e15a477505ae1b8c5676e4c21f7/libcouple/src/coupler.F90#L82>

I guess the first thing to confirm is that the `mpirun` command printed from run\_cmd has the right number of processors in the `-n` argument for MOM? The MOM exectuable is named `fms_ACCESS-OM.x` (And are the number of processors consistent with config.yaml)

---

<div class="post-metadata">

**Author:** ![angus-g](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/angus-g/32/89_2.png) [@angus-g](https://forum.access-hive.org.au/u/angus-g)\
**Post date:** [17 February 2025 03:19 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/17 "2025-02-17T03:19:25Z")

</div>

I don’t know the specifics of Leonardo, but I did think that SLURM-based systems had to use `srun` as a “wrapper” for MPI (it actually does a lot about setting up the execution environment), rather than trying to call `mpiexec` directly. That’s what we had to do on Setonix, but I’m not sure if that continues to be the approach? Apologies if this is a red herring!

---

<div class="post-metadata">

**Author:** ![Aidan](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/aidan/32/42_2.png) [@Aidan](https://forum.access-hive.org.au/u/Aidan)\
**Post date:** [17 February 2025 05:34 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/18 "2025-02-17T05:34:40Z")

</div>

I’d definitely be asking your local HPC Helpdesk @Natalia to see if they can assist with this error. I am sure they’d have seen similar issues and be able to advise.

Please also note we’re doing some exploratory work to port ACCESS models to Setonix, a SLURM based machine. I can’t give you a definite time-line, but I would think we’d have made some progress within a month.

---

<div class="post-metadata">

**Author:** ![Natalia](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/natalia/32/1612_2.png) [@Natalia](https://forum.access-hive.org.au/u/Natalia)\
**Post date:** [17 February 2025 20:41 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/19 "2025-02-17T20:41:40Z")

</div>

Not a red herring! Makes sense, tried it and got:

> The application appears to have been direct launched using “srun”,  
> but OMPI was not built with SLURM support. This usually happens  
> when OMPI was not configured --with-slurm and we weren’t able  
> to discover a SLURM installation in the usual places.
> 
> Please configure as appropriate and try again.

The thing is that I’m using a spack built OMPI from ACCESS NRI repo,  
it was likely built without SLURM support.

`which srun` points to systemwide executable even the ACCESS NRI module is loaded:

```auto
21:23 $ which srun 
/usr/bin/srun

```

Also `srun` do not support flags like `–mca io ompio --mca io\_ompio\_num\_aggregators 1, not sure whether they needed.

From what I read here [10.7. Launching with Slurm — Open MPI main documentation](https://docs.open-mpi.org/en/main/launching-apps/slurm.html)  
`mpirun` is the recommended method for launching Open MPI jobs in Slurm jobs (at least now I know 🤦‍♀️).  
Thanks!

---

<div class="post-metadata">

**Author:** ![anton](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/anton/32/1836_2.png) [@anton](https://forum.access-hive.org.au/u/anton)\
**Post date:** [17 February 2025 21:47 UTC](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914/20 "2025-02-17T21:47:27Z")

</div>

> [@Natalia](#):
>
> `–mca io ompio --mca io_ompio_num_aggregators 1`

My guess is that these settings are not essential for the model to run, and that the defaults probably work.

> The thing is that I’m using a spack built OMPI from ACCESS NRI repo,  
> it was likely built without SLURM support.

It’s possibly worth revisiting this and using the system provided OpenMPI. On gadi, we found the system provided version performed better than the version we built through spack. If there is someone at your local helpdesk who understands spack, they may have advice ?

[Next page](https://forum.access-hive.org.au/t/payu-issues-on-leonardo/3914.md?page=2)
