# UM RNS Error: ‘glm\_um\_recon1’ failed with MPI\_INIT error and missing PE0 file

**URL:** https://forum.access-hive.org.au/t/um-rns-error-glm-um-recon1-failed-with-mpi-init-error-and-missing-pe0-file/5756
**Category:** rAM3/Regional Nesting Suite
**Tags:** help
**Created:** [8 January 2026 04:17 UTC](https://forum.access-hive.org.au/t/um-rns-error-glm-um-recon1-failed-with-mpi-init-error-and-missing-pe0-file/5756 "2026-01-08T04:17:07Z")
**Posts on this page:** 10
**Page:** 1

<div class="post-metadata">

### Author: ![ZhangchengPei](https://avatars.discourse-cdn.com/v4/letter/z/8e8cbc/32.png) [@ZhangchengPei](https://forum.access-hive.org.au/u/ZhangchengPei)
#### Post date: [8 January 2026 04:17 UTC](https://forum.access-hive.org.au/t/um-rns-error-glm-um-recon1-failed-with-mpi-init-error-and-missing-pe0-file/5756/1 "2026-01-08T04:17:07Z")

</div>

Hi Atmos community,

I failed at the ‘glm\_um\_recon1’ step while running the UM RNS.

The job.err indicates:

> It looks like MPI\_INIT failed for some reason; your parallel process is  
> likely to abort. There are many reasons that a parallel process can  
> fail during MPI\_INIT; some of which are due to configuration or environment  
> problems. This failure appears to be an internal failure; here’s some  
> additional information (which may only be relevant to an Open MPI  
> developer):
> 
> ompi\_mpi\_init: ompi\_rte\_init failed
> 
> → Returned “Error” (-1) instead of “Success” (0)
> 
> \*\*\* and potentially your MPI job)  
> \*\*\* An error occurred in MPI\_Init  
> \*\*\* on a NULL communicator  
> \*\*\* MPI\_ERRORS\_ARE\_FATAL (processes in this communicator will now abort,  
> \*\*\* and potentially your MPI job)

and job.out reports:

> Could not find PE0 output file: pe\_output/umgla.fort6.pe000

I previously encountered this bug during the hh5 to xp65 transition, where reverting to an older version of `conda/analysis3` (25.05) solved it. However, that fix is no longer working.

Does anyone have suggestions on how to resolve this?

Thanks,

Zhangcheng

---

<div class="post-metadata">

### Author: ![Matt\_Woodhouse](https://avatars.discourse-cdn.com/v4/letter/m/9fc348/32.png) [@Matt\_Woodhouse](https://forum.access-hive.org.au/u/Matt_Woodhouse)
#### Post date: [8 January 2026 04:55 UTC](https://forum.access-hive.org.au/t/um-rns-error-glm-um-recon1-failed-with-mpi-init-error-and-missing-pe0-file/5756/2 "2026-01-08T04:55:33Z")

</div>

Hi Zhangcheng, I’ve been running with analysis3 24.09, which seems to be working (having previously had that error).

---

<div class="post-metadata">

### Author: ![ZhangchengPei](https://avatars.discourse-cdn.com/v4/letter/z/8e8cbc/32.png) [@ZhangchengPei](https://forum.access-hive.org.au/u/ZhangchengPei)
#### Post date: [9 January 2026 00:17 UTC](https://forum.access-hive.org.au/t/um-rns-error-glm-um-recon1-failed-with-mpi-init-error-and-missing-pe0-file/5756/3 "2026-01-09T00:17:56Z")

</div>

Hi Matt,

Thanks for the suggestions! It looks like I don’t have version 24.09 available in my conda environment—only 24.07, 24.11, and 25.\*\*. I’ve tested both 24.07 and 24.11, but unfortunately, neither resolved the issue.

Would you mind sharing your branch id? I’d like to compare our setups and see if I can spot any key differences.

Cheers,

Zhangcheng

---

<div class="post-metadata">

### Author: ![Matt\_Woodhouse](https://avatars.discourse-cdn.com/v4/letter/m/9fc348/32.png) [@Matt\_Woodhouse](https://forum.access-hive.org.au/u/Matt_Woodhouse)
#### Post date: [9 January 2026 05:06 UTC](https://forum.access-hive.org.au/t/um-rns-error-glm-um-recon1-failed-with-mpi-init-error-and-missing-pe0-file/5756/4 "2026-01-09T05:06:56Z")

</div>

Hi Zhangcheng,

I have updated to xp65 and analyis3-24.09 in a nesting suite simulation that is currently running. I’m also still using cylc7.

I’ve updated the following basis suites to match my running suite, but haven’t tested them. They also include changes to the emissions files and boundary layer nucleation options, changes which you might also like to consider.

u-df869 - glm only nesting suite to generate start dumps

u-df510 - ancillary suite

u-df403 - nesting suite

Matt

---

<div class="post-metadata">

### Author: ![Matt\_Woodhouse](https://avatars.discourse-cdn.com/v4/letter/m/9fc348/32.png) [@Matt\_Woodhouse](https://forum.access-hive.org.au/u/Matt_Woodhouse)
#### Post date: [9 January 2026 06:04 UTC](https://forum.access-hive.org.au/t/um-rns-error-glm-um-recon1-failed-with-mpi-init-error-and-missing-pe0-file/5756/5 "2026-01-09T06:04:05Z")

</div>

I think I meant 24.11, not .09

Though I now seem to be getting the error, having just had a suite complete successfully.

---

<div class="post-metadata">

### Author: ![ZhangchengPei](https://avatars.discourse-cdn.com/v4/letter/z/8e8cbc/32.png) [@ZhangchengPei](https://forum.access-hive.org.au/u/ZhangchengPei)
#### Post date: [11 January 2026 23:20 UTC](https://forum.access-hive.org.au/t/um-rns-error-glm-um-recon1-failed-with-mpi-init-error-and-missing-pe0-file/5756/6 "2026-01-11T23:20:19Z")

</div>

It seems like something in the environment has changed, specifically regarding the openmpi. Is anyone familiar with this issue?

---

<div class="post-metadata">

### Author: ![lachlanswhyborn](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/lachlanswhyborn/32/1837_2.png) [@lachlanswhyborn](https://forum.access-hive.org.au/u/lachlanswhyborn)
#### Post date: [12 January 2026 03:06 UTC](https://forum.access-hive.org.au/t/um-rns-error-glm-um-recon1-failed-with-mpi-init-error-and-missing-pe0-file/5756/7 "2026-01-12T03:06:24Z")

</div>

Is there a reason you’re loading the `conda/analysis` module to run the RNS?

---

<div class="post-metadata">

### Author: ![ZhangchengPei](https://avatars.discourse-cdn.com/v4/letter/z/8e8cbc/32.png) [@ZhangchengPei](https://forum.access-hive.org.au/u/ZhangchengPei)
#### Post date: [12 January 2026 05:11 UTC](https://forum.access-hive.org.au/t/um-rns-error-glm-um-recon1-failed-with-mpi-init-error-and-missing-pe0-file/5756/8 "2026-01-12T05:11:51Z")

</div>

Hi Lachlan,

It appears the RNS requires the ‘pytz’ module to enable model cycling, as Bec has mentioned in this post [Using xp65 in UM suites](https://forum.access-hive.org.au/t/using-xp65-in-um-suites/5292). I tested this by not loading the conda/analysis environment, which resulted in a

> ModuleNotFoundError: No module named ‘pytz’.

---

<div class="post-metadata">

### Author: ![lachlanswhyborn](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/lachlanswhyborn/32/1837_2.png) [@lachlanswhyborn](https://forum.access-hive.org.au/u/lachlanswhyborn)
#### Post date: [12 January 2026 05:24 UTC](https://forum.access-hive.org.au/t/um-rns-error-glm-um-recon1-failed-with-mpi-init-error-and-missing-pe0-file/5756/9 "2026-01-12T05:24:13Z")

</div>

The xp65 Conda environment overrides the `openmpi` which is loaded with `module load openmpi/x.y.z` (should be able to see this with `echo $OPAL_PREFIX` with the `conda/analysis` module loaded). You might be able to get around this by unsetting `$OPAL_PREFIX` (i.e. setting it to an empty string) after loading the Conda environment, but I’m not sure.

---

<div class="post-metadata">

### Author: ![ZhangchengPei](https://avatars.discourse-cdn.com/v4/letter/z/8e8cbc/32.png) [@ZhangchengPei](https://forum.access-hive.org.au/u/ZhangchengPei)
#### Post date: [12 January 2026 23:46 UTC](https://forum.access-hive.org.au/t/um-rns-error-glm-um-recon1-failed-with-mpi-init-error-and-missing-pe0-file/5756/10 "2026-01-12T23:46:55Z")

</div>

Update: I load python3 instead of the default python2 in PRE\_COMMAND and removed the conda/analysis. The RNS is now running successfully! Thank you for the suggestions @Matt_Woodhouse and @lachlanswhyborn .
