# Very different zarr file sizes from virtually identical write operations

**URL:** <https://forum.access-hive.org.au/t/very-different-zarr-file-sizes-from-virtually-identical-write-operations/955>\
**Category:** Technical\
**Tags:** zarr, python\
**Created:** [6 July 2023 00:16 UTC](https://forum.access-hive.org.au/t/very-different-zarr-file-sizes-from-virtually-identical-write-operations/955 "2023-07-06T00:16:58Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![dougrichardson](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/dougrichardson/32/523_2.png) [@dougrichardson](https://forum.access-hive.org.au/u/dougrichardson)\
**Post date:** [6 July 2023 00:16 UTC](https://forum.access-hive.org.au/t/very-different-zarr-file-sizes-from-virtually-identical-write-operations/955/1 "2023-07-06T00:16:58Z")

</div>

I have computed two nearly identical metrics from daily ERA5 temperature data, given by the below functions `calc_cdd` and `calc_hdd`. As you can see, the only difference is the direction of the inequality in `.where()` and the order of the subtraction.

```auto
def calc_cdd(T, comfort=24):
    return (T - comfort).where(T > comfort, 0)

def calc_hdd(T, comfort=18):
    return (comfort - T).where(T < comfort, 0)

```

I spin up a full node using `PBScluster`, load my temperature data (~90 GB) with what I think is sensible chunking, apply the functions and a few other simple operations, and chunk again before writing to zarr:

```auto
T = xr.open_mfdataset(era_path+"2t/daily/*.nc", chunks={"time": "200MB"})

T = T - 273.15

```

 ![image](https://us1.discourse-cdn.com/flex020/uploads/access1/original/1X/93a5da608c119b035b17d00801a946fe2c86c3ee.png)

```auto
cdd = calc_cdd(T)
cdd = cdd.rename({"t2m": "cdd"})

# Need to chunk again so that we have uniform chunk sizes
cdd = cdd.chunk({"time": "200MB"})

cdd.to_zarr(
    era_path + "/derived/cdd_24_era5_daily_1959-2022.zarr",
    mode="w",
    consolidated=True
)

```

The computation for `calc_hdd` takes a bit longer than for `calc_cdd`. More concerningly, the resultant `zarr` files have very different sizes, with the former being over double the size:

 ![image](https://us1.discourse-cdn.com/flex020/uploads/access1/original/1X/cf15a333443609267a06852f5d69488db9ad2292.png)

What are the possible reasons for this? It doesn’t necessarily matter for my workflow, but if there’s a way to keep them both as small as possible that would be convenient.

---

<div class="post-metadata">

**Author:** ![dougiesquire](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/dougiesquire/32/16_2.png) [@dougiesquire](https://forum.access-hive.org.au/u/dougiesquire)\
**Post date:** [6 July 2023 01:51 UTC](https://forum.access-hive.org.au/t/very-different-zarr-file-sizes-from-virtually-identical-write-operations/955/2 "2023-07-06T01:51:25Z")

</div>

Hi @dougrichardson. I expect the difference in size is just due to the compression applied by zarr. Are the uncompressed sizes the same if you read your zarr collections back in using xarray (e.g. see the size listed in the Array column and Bytes row of the DataArray html repr)?

It’s possibly also worth checking that all the zarr chunks were successfully written. With `consolidate=True` it’s possible to create a zarr collection that appears fine when you first open it lazily (from the consolidated metadata), but actually has data missing (e.g. if the writing of the chunks failed at some point, or if some chunks were deleted after the collection was created). For peace of mind, you could make sure that there are no NaNs in your zarr collections.

---

<div class="post-metadata">

**Author:** ![dougrichardson](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/dougrichardson/32/523_2.png) [@dougrichardson](https://forum.access-hive.org.au/u/dougrichardson)\
**Post date:** [6 July 2023 02:00 UTC](https://forum.access-hive.org.au/t/very-different-zarr-file-sizes-from-virtually-identical-write-operations/955/3 "2023-07-06T02:00:45Z")

</div>

Thanks @dougiesquire. The uncompressed sizes are the same, and there are no NaNs, so compression seems to be the answer. It doesn’t make sense to me given the two operations are nearly the same, but also doesn’t really matter.

---

<div class="post-metadata">

**Author:** ![dougiesquire](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/dougiesquire/32/16_2.png) [@dougiesquire](https://forum.access-hive.org.au/u/dougiesquire)\
**Post date:** [6 July 2023 02:27 UTC](https://forum.access-hive.org.au/t/very-different-zarr-file-sizes-from-virtually-identical-write-operations/955/4 "2023-07-06T02:27:01Z")

</div>

File compression algorithms are complex, but in general they try to get rid of redundancy (I think zarr uses blosc by default, if you want to read about it). So if there are lots of the same or similar values in a file, it will be easier to get a high compression ratio.

---

<div class="post-metadata">

**Author:** ![Aidan](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/aidan/32/42_2.png) [@Aidan](https://forum.access-hive.org.au/u/Aidan)\
**Post date:** [6 July 2023 05:49 UTC](https://forum.access-hive.org.au/t/very-different-zarr-file-sizes-from-virtually-identical-write-operations/955/5 "2023-07-06T05:49:22Z")

</div>

> [@dougiesquire](#):
>
> So if there are lots of the same or similar values in a file, it will be easier to get a high compression ratio

I might expect there to be a big difference if you were masking out large contiguous parts of the output and the sizes of the masked area was markedly different in the two cases.

You’re masking to zero, do the sizes of those masked areas differ a lot in the two cases?

---

<div class="post-metadata">

**Author:** ![dougrichardson](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/dougrichardson/32/523_2.png) [@dougrichardson](https://forum.access-hive.org.au/u/dougrichardson)\
**Post date:** [7 July 2023 03:46 UTC](https://forum.access-hive.org.au/t/very-different-zarr-file-sizes-from-virtually-identical-write-operations/955/6 "2023-07-07T03:46:24Z")

</div>

Hi @Aidan, I just checked this and there are many more zeros in the smaller `zarr` store.

---

<div class="post-metadata">

**Author:** ![Aidan](https://sea2.discourse-cdn.com/flex020/user_avatar/forum.access-hive.org.au/aidan/32/42_2.png) [@Aidan](https://forum.access-hive.org.au/u/Aidan)\
**Post date:** [7 July 2023 05:14 UTC](https://forum.access-hive.org.au/t/very-different-zarr-file-sizes-from-virtually-identical-write-operations/955/7 "2023-07-07T05:14:36Z")

</div>

> [@dougrichardson](#):
>
> I just checked this and there are many more zeros in the smaller `zarr` store.

Right, so that is consistent due to the compressibility of essentially zero information.
