Skip to content

UxDataset.to_xarray() and UxDataArray.to_xarray() should use _construct_direct #1794

Description

@Sevans711

Converting uxarray objects to xarray shouldn't require running all of xarray's validation checks during xr.DataArray.__init__ and xr.Dataset.__init__, because every UxDataArray / UxDataset is already a valid xr.DataArray / xr.Dataset. These checks scale with number of coordinates and variables, and they don't take extremely long so this probably isn't a huge priority, but the overhead could become noticeable in any hot loop handling bookkeeping tasks like this.

For example, in some local tests I swapped UxDataset.to_xarray() to xr.Dataset._construct_direct, and measured speedups (via %%timeit) of:

UxDataset size main improved
100 data_vars, 2.8 GB, 150k faces 26 ms ± 169 μs 4.65 μs ± 31.1 ns
265 data_vars, 4.1 GB, 150k faces 118 ms ± 614 μs 5.27 μs ± 14.4 ns

The new UxDataset.to_xarray() implementation was like this:

    def to_xarray(self, grid_format: str = "UGRID") -> xr.Dataset:
        if grid_format == "HEALPix":
                ds = self.rename_dims({"n_face": "cell"})
        else:
            ds = self

        return xr.Dataset._construct_direct(
            dict(ds._variables),
            set(ds._coord_names),
            dims=dict(ds._dims),
            attrs=dict(ds.attrs),
            indexes=dict(ds._indexes),
        )

(Claude suggested this implementation, but I read/understand it, and ran the timing tests myself.)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

scalabilityRelated to scalability & performance efforts

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions