# Need help finding and interpreting ACCESS-ESM1-5 errors

**URL:** <https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300>\
**Category:** Earth System\
**Tags:** help\
**Created:** [19 August 2024 23:25 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300 "2024-08-19T23:25:20Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![hrsdawson](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/hrsdawson/32/218_2.png) [@hrsdawson](https://forum.access-hive.org.au/u/hrsdawson)\
**Post date:** [19 August 2024 23:25 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/1 "2024-08-19T23:25:20Z")

</div>

I’m running some perturbation experiments with the pre-industrial control in ACCESS-ESM1-5 but since last week a number of them have started crashing during the run and I don’t know why.

I get the error message: “_payu: Model exited with error code 139; aborting._”. I can see this has been mentioned in [other posts,](https://forum.access-hive.org.au/t/run-access-esm-fails-with-error-code-139/1749) but I don’t think this is relevant here because I have started these experiments using the PI-02 restarts.

I’d like some help finding and interpreting the errors. An example control directory which has crashed at time step 4525 is here: `/home/561/hd4873/PostDoc/ACCESS-ESM/access-esm-payu`

The work directory is here: `/scratch/e14/hd4873/access-esm/work/access-esm-payu-ocean-warm-upwelling_year780-6213760d`

and the error logs are here: `/scratch/e14/hd4873/access-esm/archive/access-esm-payu-ocean-warm-upwelling_year780-6213760d`.

---

<div class="post-metadata">

**Author:** ![hrsdawson](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/hrsdawson/32/218_2.png) [@hrsdawson](https://forum.access-hive.org.au/u/hrsdawson)\
**Post date:** [20 August 2024 00:19 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/2 "2024-08-20T00:19:22Z")

</div>

I think it crashed in the UM but not sure which atmosphere logs give more details. This is a snippet from the access.err file.

```auto
[gadi-cpu-clx-0429:974522:0:974522] Caught signal 11 (Segmentation fault: Sent by the kernel at address (nil))
--------------------------------------------------------------------------
Primary job terminated normally, but 1 process returned
a non-zero exit code. Per user-direction, the job has been aborted.
--------------------------------------------------------------------------
forrtl: error (78): process killed (SIGTERM)
Image PC Routine Line Source
um7.3x 00000000012FCBC4 Unknown Unknown Unknown
libpthread-2.28.s 0000154A43EB5D20 Unknown Unknown Unknown
mca_pml_ucx.so 0000154A30389E27 mca_pml_ucx_recv Unknown Unknown
libmpi.so.40.20.2 0000154A444CAF55 MPI_Recv Unknown Unknown
libmpi_mpifh.so 0000154A447CA170 Unknown Unknown Unknown
um7.3x 0000000001111A08 mpl_recv_ 67 mpl_recv.F90
um7.3x 000000000110AF66 gc_rrecv_ 168 gc_rrecv.F90
um7.3x 0000000000989C5D bi_linear_h_ 613 bi_linear_h.f90
um7.3x 0000000000CC92BC ritchie_ 2557 ritchie.f90
um7.3x 00000000009DFB7D departure_point_ 382 departure_point.f90
um7.3x 00000000008AA46F sl_thermo_ 681 sl_thermo.f90
um7.3x 00000000006EC376 ni_sl_thermo_ 778 ni_sl_thermo.f90
um7.3x 00000000004BD235 Unknown Unknown Unknown
um7.3x 0000000000435BB0 Unknown Unknown Unknown
um7.3x 000000000041481E um_shell_ 3930 um_shell.f90
um7.3x 000000000040D968 MAIN__ 40 flumeMain.f90
um7.3x 000000000040D8A2 Unknown Unknown Unknown
libc-2.28.so 0000154A439037E5 __libc_start_main Unknown Unknown
um7.3x 000000000040D7AE Unknown Unknown Unknown

```

---

<div class="post-metadata">

**Author:** ![dkhutch](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/dkhutch/32/510_2.png) [@dkhutch](https://forum.access-hive.org.au/u/dkhutch)\
**Post date:** [20 August 2024 00:52 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/3 "2024-08-20T00:52:44Z")

</div>

I know it might sound unlikely, but have you checked if the error occurs again if you simply do a sweep and re-run it?  
ACCESS-ESM1.5 has a habit of crashing for unknown reasons, and I find it’s best to first check if the error occurs twice. (Sometimes it just runs fine the second time you try…)

---

<div class="post-metadata">

**Author:** ![hrsdawson](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/hrsdawson/32/218_2.png) [@hrsdawson](https://forum.access-hive.org.au/u/hrsdawson)\
**Post date:** [20 August 2024 01:40 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/4 "2024-08-20T01:40:58Z")

</div>

Yep, I’ve tried that. If I recall correctly, it’s crashed at the exact same time step. I can try again now to confirm.

---

<div class="post-metadata">

**Author:** ![dkhutch](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/dkhutch/32/510_2.png) [@dkhutch](https://forum.access-hive.org.au/u/dkhutch)\
**Post date:** [20 August 2024 01:58 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/5 "2024-08-20T01:58:36Z")

</div>

Twice is enough to confirm. Thanks Hannah.

Dr David Hutchinson (he/him)  
ARC DECRA fellow in paleoclimate modelling  
Climate Change Research Centre, UNSW Sydney  
david.hutchinson@unsw.edu.au

---

<div class="post-metadata">

**Author:** ![hrsdawson](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/hrsdawson/32/218_2.png) [@hrsdawson](https://forum.access-hive.org.au/u/hrsdawson)\
**Post date:** [20 August 2024 02:14 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/6 "2024-08-20T02:14:57Z")

</div>

yep, confirming it’s crashed at the same spot.

---

<div class="post-metadata">

**Author:** ![HIMADRI\_SAINI](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/himadri_saini/32/515_2.png) [@HIMADRI\_SAINI](https://forum.access-hive.org.au/u/HIMADRI_SAINI)\
**Post date:** [20 August 2024 02:22 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/7 "2024-08-20T02:22:37Z")

</div>

Have you tried perturbing the model using /projects/access/apps/pythonlib/umfile\_utils/perturbIC.py. I encountered this error, and I’m not sure why, but perturbing sometimes does the trick.

---

<div class="post-metadata">

**Author:** ![Aidan](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/aidan/32/42_2.png) [@Aidan](https://forum.access-hive.org.au/u/Aidan)\
**Post date:** [20 August 2024 02:30 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/8 "2024-08-20T02:30:05Z")

</div>

> [@hrsdawson](#):
>
> ` 0000000000989C5D bi_linear_h_ 613 bi_linear_h.f90`

A crash in the `bi_linear_h` routine is often the result of the model becoming unstable and if so can fixed with a small perturbation of the atmosphere as @HIMADRI_SAINI suggested.

See also

[http://climate-cms.wikis.unsw.edu.au/ACCESS#Coupled\_Model\_Crashes](http://climate-cms.wikis.unsw.edu.au/ACCESS#Coupled_Model_Crashes)

---

<div class="post-metadata">

**Author:** ![hrsdawson](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/hrsdawson/32/218_2.png) [@hrsdawson](https://forum.access-hive.org.au/u/hrsdawson)\
**Post date:** [20 August 2024 03:51 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/9 "2024-08-20T03:51:52Z")

</div>

Thanks for the suggestion @HIMADRI_SAINI. I’ve not tried this. I’ll give it a go.

---

<div class="post-metadata">

**Author:** ![hrsdawson](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/hrsdawson/32/218_2.png) [@hrsdawson](https://forum.access-hive.org.au/u/hrsdawson)\
**Post date:** [20 August 2024 04:14 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/10 "2024-08-20T04:14:25Z")

</div>

@Aidan @HIMADRI_SAINI sorry, what’s the syntax for running this? I’ve just tried like so (from the CLEX instructions): `/projects/access/apps/pythonlib/umfile_utils/perturbIC.py restart_dump.astart`

and I get the following message:

```auto
  File "/projects/access/apps/pythonlib/umfile_utils/perturbIC.py", line 23
    print "Usage: perturbIC [-a amplitude] [-v variable (stashcode)] file"
          ^

```

So I tried copying the script and running `perturbIC -arguments file` as suggested above (what’s considered a small perturbation by the way?), but no luck there - think I’m getting the syntax wrong.

---

<div class="post-metadata">

**Author:** ![dkhutch](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/dkhutch/32/510_2.png) [@dkhutch](https://forum.access-hive.org.au/u/dkhutch)\
**Post date:** [20 August 2024 04:37 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/11 "2024-08-20T04:37:44Z")

</div>

Hi Hannah,  
This seems like a python2 / python3 problem. The perturbIC.py in the example is using the old print method, where print is done without brackets. A couple of ways to deal with this (not sure what Himadri did…):

- explicitly call python2 interpreter
- copy the script to a local directory and update to print() with brackets.

Probably the first way is easier because the script has dependencies to the umfile.py and um\_fileheaders.py script in the same directory.

---

<div class="post-metadata">

**Author:** ![dkhutch](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/dkhutch/32/510_2.png) [@dkhutch](https://forum.access-hive.org.au/u/dkhutch)\
**Post date:** [20 August 2024 04:41 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/12 "2024-08-20T04:41:01Z")

</div>

By default, the perturbIC script makes a perturbation of order 0.01 to the temperature field.

---

<div class="post-metadata">

**Author:** ![hrsdawson](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/hrsdawson/32/218_2.png) [@hrsdawson](https://forum.access-hive.org.au/u/hrsdawson)\
**Post date:** [20 August 2024 04:50 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/13 "2024-08-20T04:50:25Z")

</div>

@dkhutch thanks! I loaded python2/2.1.7 and ran `python2 /projects/access/apps/pythonlib/umfile_utils/perturbIC.py restart_dump.astart` and I think that’s worked.

---

<div class="post-metadata">

**Author:** ![MartinDix](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/martindix/32/32_2.png) [@MartinDix](https://forum.access-hive.org.au/u/MartinDix)\
**Post date:** [20 August 2024 05:40 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/14 "2024-08-20T05:40:28Z")

</div>

There’s also a version updated for python3 available in the `pythonlib/umfile_utils/access_cm2` module. It differs in requiring a seed as an argument which makes it reproducible

```auto
module use ~access/modules
module load pythonlib/umfile_utils/access_cm2 
perturbIC.py -h
usage: perturbIC.py [-h] [-a AMPLITUDE] -s SEED ifile

Perturb UM initial dump

positional arguments:
  ifile Input file (modified in place)

options:
  -h, --help show this help message and exit
  -a AMPLITUDE Amplitude of perturbation
  -s SEED Random number seed (must be non-negative integer)

```

---

<div class="post-metadata">

**Author:** ![hrsdawson](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/hrsdawson/32/218_2.png) [@hrsdawson](https://forum.access-hive.org.au/u/hrsdawson)\
**Post date:** [20 August 2024 06:44 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/15 "2024-08-20T06:44:58Z")

</div>

Great, that worked - thanks all for your help!

What does this mean for publication purposes though? At the moment I’m testing so it’s fine, but if this had to be done for an ensemble that I wanted to publish - is this perturbing method considered okay by the Earth System community if you use the reproducible method that @MartinDix posted? Or is it not really okay to publish runs that have had to be perturbed in this way?

---

<div class="post-metadata">

**Author:** ![dkhutch](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/dkhutch/32/510_2.png) [@dkhutch](https://forum.access-hive.org.au/u/dkhutch)\
**Post date:** [20 August 2024 08:03 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/16 "2024-08-20T08:03:10Z")

</div>

My feeling is that making a small perturbation during spin up like this is not at all problematic for publication. Plenty of models have to do tricks like this… It is basically a non-issue. (If you want to document instances of perturbations like this then great, that’s probably better than what most people do.)

Dr David Hutchinson (he/him)  
ARC DECRA fellow in paleoclimate modelling  
Climate Change Research Centre, UNSW Sydney  
david.hutchinson@unsw.edu.au

---

<div class="post-metadata">

**Author:** ![Aidan](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/aidan/32/42_2.png) [@Aidan](https://forum.access-hive.org.au/u/Aidan)\
**Post date:** [20 August 2024 09:41 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/17 "2024-08-20T09:41:29Z")

</div>

This was covered in a previous topic

> [@How to preserve reproducibility when applying perturbations in the event of a numerical instability crash](https://forum.access-hive.org.au/t/how-to-preserve-reproducibility-when-applying-perturbations-in-the-event-of-a-numerical-instability-crash/411/8):
>
> I can’t see any issue with what you’ve suggested @holger. To be explicit, I would do as @holger suggests, with these specific steps Run the perturbIC.pywith known seed Then payu setup, which will rewrite the manifest file with your new (perturbed) restart(s) git commit -a and write a commit message documenting the steps you have taken to perturb the restarts, with the seed value and the location of the script used payu run You shouldn’t need to invoke a specific script with userscripts. If …

---

<div class="post-metadata">

**Author:** ![hrsdawson](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/hrsdawson/32/218_2.png) [@hrsdawson](https://forum.access-hive.org.au/u/hrsdawson)\
**Post date:** [20 August 2024 23:37 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/18 "2024-08-20T23:37:56Z")

</div>

Oops I hadn’t seen that, thank you for linking @Aidan.

---

<div class="post-metadata">

**Author:** ![Aidan](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/aidan/32/42_2.png) [@Aidan](https://forum.access-hive.org.au/u/Aidan)\
**Post date:** [21 August 2024 00:25 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/19 "2024-08-21T00:25:21Z")

</div>

No worries. Hope it is helpful.

---

<div class="post-metadata">

**Author:** ![anton](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/anton/32/1836_2.png) [@anton](https://forum.access-hive.org.au/u/anton)\
**Post date:** [27 August 2024 23:25 UTC](https://forum.access-hive.org.au/t/need-help-finding-and-interpreting-access-esm1-5-errors/2300/20 "2024-08-27T23:25:26Z")

</div>


