Converting uxarray objects to xarray shouldn't require running all of xarray's validation checks during xr.DataArray.__init__ and xr.Dataset.__init__, because every UxDataArray / UxDataset is already a valid xr.DataArray / xr.Dataset. These checks scale with number of coordinates and variables, and they don't take extremely long so this probably isn't a huge priority, but the overhead could become noticeable in any hot loop handling bookkeeping tasks like this.
For example, in some local tests I swapped UxDataset.to_xarray() to xr.Dataset._construct_direct, and measured speedups (via %%timeit) of:
| UxDataset size |
main |
improved |
| 100 data_vars, 2.8 GB, 150k faces |
26 ms ± 169 μs |
4.65 μs ± 31.1 ns |
| 265 data_vars, 4.1 GB, 150k faces |
118 ms ± 614 μs |
5.27 μs ± 14.4 ns |
The new UxDataset.to_xarray() implementation was like this:
def to_xarray(self, grid_format: str = "UGRID") -> xr.Dataset:
if grid_format == "HEALPix":
ds = self.rename_dims({"n_face": "cell"})
else:
ds = self
return xr.Dataset._construct_direct(
dict(ds._variables),
set(ds._coord_names),
dims=dict(ds._dims),
attrs=dict(ds.attrs),
indexes=dict(ds._indexes),
)
(Claude suggested this implementation, but I read/understand it, and ran the timing tests myself.)
Converting uxarray objects to xarray shouldn't require running all of xarray's validation checks during
xr.DataArray.__init__andxr.Dataset.__init__, because every UxDataArray / UxDataset is already a valid xr.DataArray / xr.Dataset. These checks scale with number of coordinates and variables, and they don't take extremely long so this probably isn't a huge priority, but the overhead could become noticeable in any hot loop handling bookkeeping tasks like this.For example, in some local tests I swapped UxDataset.to_xarray() to
xr.Dataset._construct_direct, and measured speedups (via%%timeit) of:The new UxDataset.to_xarray() implementation was like this:
(Claude suggested this implementation, but I read/understand it, and ran the timing tests myself.)