Skip to content

enh(data): add page on data to the tests section - #580

Open
lwasser wants to merge 2 commits into
pyOpenSci:mainfrom
lwasser:data-david
Open

lwasser wants to merge 2 commits into
pyOpenSci:mainfrom
lwasser:data-david

Conversation

@lwasser

@lwasser lwasser commented Sep 19, 2025

Copy link
Copy Markdown
Member

This pr is a rework of #110 . Let's plan to run a sprint on this pr in the next few weeks to see if we can get to a place where it's good.

NOTE: I added some code examples that I literally found online and DID NOT TEST. so we will want to definitely test what is there before merging.

ALSO - because this topic is not my expertise, in some cases I ran with a section after a Google search and tried to flesh it out, but it also could be off. Any feedback is welcome on this!!

@lwasser lwasser added help wanted We welcome a contributor to work on this issue! thank you in advance! 🚀 ready-for-review labels Sep 19, 2025
Comment thread tests/package-data.md
* **[Open Science Framework (OSF)](https://osf.io)** - Comprehensive research platform
* **[Figshare](https://figshare.com)** - User-friendly with good visualization tools
* **[Dryad](https://datadryad.org)** - Focused on research data (subscription model for some features)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Would huggingface fit here or is this more for academic ? Perhaps https://docs.source.coop ?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's good you suggested these -- sorry for a big wordy response to a single suggestion, but I think it shows that we could say a little more about why we have what we have here.

I'll say what I think we might add, but first, re: your suggestions:
Huggingface is free, and convenient.
I have a couple of worries about it though. It will only exist (1) at most as long as the company exists, and that might not be long if the AI bubble pops, and (2) it might not even be that long if the company decides that the cost of maintaining all the data is not worth how many customers it gets them. I don't think this is just an academic thing, since a company could choose to be structured in such a way that they can't pull the rug on users whenever they want. Forgive me for being pedantic about it, I guess I'm sensitive as a failed academic 😇

Source Co-op looks cool! Maybe we should include them! Need to read about it more

re: the places we have listed now, I think there's a couple things we could do here:

  • say a little bit more in the blurb about what our criteria are. E.g., is it free, is there some sort of guarantee on how long the data will be stored, is there a good UI / API / existing software tool that makes it easy to work with
  • put in some sort of table for each of these that indicates how each place to store data stack up WRT to those criteria

Again, sorry for being wordy -- some of this stuff I say already in the actual test, but you're making me see how this could maybe be better organized

Comment thread tests/package-data.md
* **[Google Cloud Platform](https://cloud.google.com/storage)** - Cloud Storage with strong AI/ML tool integration
* **[Linode](https://www.linode.com/products/object-storage/)** - Object Storage with straightforward pricing and developer-friendly tools

<!-- I don't understand how these platforms are different from things like figshare - can we clarify that? and how / why someone would pick these vs figshare / dryad? -->

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This section is moving into data versioning which is important for scientific studies and reproducibility, but may be out of scope for a package. Having a small subset or testable data for a package makes a lot of sense.

Comment thread tests/package-data.md
- **[Pachyderm](https://www.pachyderm.com/)** - Data pipeline platform with version control
```

<!-- I am not sure how this section relates to data stored in a package. I understand it's important, but does it belong in a page focused on how and where to store data for your Python package? It might be that I just don't understand as written!-->

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1 to removing

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I hear you -- I put this in thinking "I wish I knew about this earlier so we should mention it" but it is probably out of scope for the specifics of including data with your package

Comment thread tests/package-data.md

### Use Pytest fixtures for data access

Pytest fixtures provide a clean way to set up and share data across your test suite. They're especially useful for scientific packages where you need consistent access to test datasets.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe this is just personal preference, but I think we should move pytest up and pooch down.
Pytest fits better with the whole 'testing your data thing' along with testing your code. Pooch seems like an additional tool or service to download data from an external source, I've not heard of it before.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👍 We definitely should talk about pytest fixtures sooner if this page is about using data in tests.

@lwasser (at the risk of being even more annoying on this review 😅 ) looking at this again, I'm wondering if it would make sense to break this into two pages.

One on "data for tests" and another on "example data in your package", that would not live in the section on tests, but instead live in ... some section that doesn't exist yet.
I guess it would be some section that's "packaging", but more about, like library structure I guess? As oppossed to the nitty-gritty of publishing a package

I know a lot of times test data and example data overlap but it's striking me as odd to have content about example data in the tests/ section

ok I'll stop being noisy on the review 🤐

@NickleDave NickleDave Oct 7, 2025 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One last thought: if I was writing a page that was only "data for tests", I'd write something like:

  1. as much as possible, use pytest.fixture with "fake data" you generate to test: this lets you avoid adding real data to your project, and helps you avoid writing code and tests that are tightly coupled, encouraging you to test the interface instead of implementation details.
  2. for functionality where you absolutely need to test on real data, e.g. parsing specific file formats, include those in version control if possible, but avoid pushing the data to PyPI because it increases package size and strains the service
  3. in some cases you may include example data in your package (link to section on example data), and you can re-use this for tests, using fixtures, here's how...
  4. to test some functionality, e.g. fitting statistical models, you may not be able to avoid downloading relatively large amounts of data for your tests. In these cases you may find it convenient to publish your dataset to open data repositories and then download it as part of a set-up step for your project, that you should outline in your contributing.md. Tools exist for accessing these dataset [link to example data page again]

Comment thread tests/package-data.md
<!-- I am not sure about this statement in terms of what it means and whether we have tools that consider standards or not we might also want to link to FAIR-->
:::

```{admonition} Field specific standards + metadata

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A link to FAIR would be good. I haven't heard of the ones mentioned 😓

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1 for saying that in general datasets should be FAIR.

I could add link to the neuro formats.

I know some of the astro ones too, e.g., FITS, and the loftily-named [Advanced Scientific Data Format](https://proceedings.scipy.org/articles/majora-212e5952-000. Wwe could ask astro people.
Would geo formats be worth a mention here too?
Or is all that info overload

@lwasser

lwasser commented Nov 12, 2025

Copy link
Copy Markdown
Member Author

YAY! ok there is feedback here - y'all i'm catchin up after travel and a workshop. I'll focus on incorporating everyone's comments and then we can do another round of review.

Comment thread tests/package-data.md

Here, you will learn about working with data for your scientific Python package.

::{admonition} What you will learn

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The "What you will learn" box is not rendering as a directive. Let's fix that.

Suggested change
::{admonition} What you will learn
:::{admonition} What you will learn

Comment thread tests/package-data.md

* **movingpandas** is a pyOpenSci-accepted package that uses data to support some of its tutorials. [Here is an example of a tutorial that shows how to use MovingPandas to process bird migration data.](https://movingpandas.github.io/movingpandas-website/2-analysis-examples/bird-migration.html)

* [scikit-image:](https://github.com/scikit-image/scikit-image/tree/main/skimage/data) stores data within the package itself to be used for package examples. <i made this up>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
* [scikit-image:](https://github.com/scikit-image/scikit-image/tree/main/skimage/data) stores data within the package itself to be used for package examples. <i made this up>
* [scikit-image](https://github.com/scikit-image/scikit-image/tree/main/skimage/data)

We could expand in a follow-up PR.

Comment thread tests/package-data.md
* core scientific-python packages
* scikit-learn: <https://github.com/scikit-learn/scikit-learn/tree/main/sklearn/datasets/data>

# Store Your Data in a Scientific Repository

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

H1 breaks the sidebar nav, and I'm not sure it's needed. Consider using H2 or H3.
Placing under "Where to store your data" with H3 seems reasonable.

Comment thread tests/package-data.md
When creating data for your package, be aware of field-specific standards and formats. For example, in neuroscience, DANDI and NWB are common file formats used in the domain. So consider whether your package can / should support those formats if you expect users in the neuroscience space to use it.

???Many pyOpenSci tools exist to address these standards or to provide interoperability because these standards don't exist ???
see also: FAIR data

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
see also: FAIR data
Where possible, follow [FAIR data principles](https://www.go-fair.org/fair-principles/) so your example and test data are findable, accessible, interoperable, and reusable.

Comment thread tests/package-data.md

When creating data for your package, be aware of field-specific standards and formats. For example, in neuroscience, DANDI and NWB are common file formats used in the domain. So consider whether your package can / should support those formats if you expect users in the neuroscience space to use it.

???Many pyOpenSci tools exist to address these standards or to provide interoperability because these standards don't exist ???

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What about:

Suggested change
???Many pyOpenSci tools exist to address these standards or to provide interoperability because these standards don't exist ???
Many pyOpenSci tools are built to support these field-specific standards, or to provide interoperability between formats where no common standard exists.

Comment thread tests/package-data.md
from importlib import resources
import my_package

with resources.open_text(my_package.data, 'example.csv') as f:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

open_text is deprecated since Python 3.11. files() is the current API that works on 3.9+ and in the backport.

Suggested change
with resources.open_text(my_package.data, 'example.csv') as f:
with resources.files("my_package.data").joinpath("example.csv").open("r") as f:

Comment thread tests/package-data.md
Comment on lines +296 to +302
def load_sample_data():
file_path = data_registry.fetch("sample_data.csv")
return pd.read_csv(file_path)

def load_large_dataset():
file_path = data_registry.fetch("large_dataset.nc")
return xr.open_dataset(file_path)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixes mixed indentation, will raise an error otherwise.

Suggested change
def load_sample_data():
file_path = data_registry.fetch("sample_data.csv")
return pd.read_csv(file_path)
def load_large_dataset():
file_path = data_registry.fetch("large_dataset.nc")
return xr.open_dataset(file_path)
def load_sample_data():
file_path = data_registry.fetch("sample_data.csv")
return pd.read_csv(file_path)
def load_large_dataset():
file_path = data_registry.fetch("large_dataset.nc")
return xr.open_dataset(file_path)

Comment thread tests/package-data.md
Comment on lines +319 to +328
@pytest.fixture
def sample_data():
"""Load sample dataset for testing."""
data_path = Path(__file__).parent / "data" / "sample.csv"
return pd.read_csv(data_path)

def test_data_processing(sample_data):
"""Test uses the fixture automatically."""
result = my_function(sample_data)
assert len(result) > 0

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mixed indentation as above.

Suggested change
@pytest.fixture
def sample_data():
"""Load sample dataset for testing."""
data_path = Path(__file__).parent / "data" / "sample.csv"
return pd.read_csv(data_path)
def test_data_processing(sample_data):
"""Test uses the fixture automatically."""
result = my_function(sample_data)
assert len(result) > 0
@pytest.fixture
def sample_data():
"""Load sample dataset for testing."""
data_path = Path(__file__).parent / "data" / "sample.csv"
return pd.read_csv(data_path)
def test_data_processing(sample_data):
"""Test uses the fixture automatically."""
result = my_function(sample_data)
assert len(result) > 0

Comment thread tests/package-data.md
Comment on lines +339 to +348
@pytest.fixture(scope="session")
def remote_dataset():
"""Download and cache remote data once per test session."""
file_path = data_registry.fetch("sample_data.csv")
return pd.read_csv(file_path)

def test_remote_data_analysis(remote_dataset):
"""Test using remote dataset."""
result = analyze_dataset(remote_dataset)
assert result is not None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same indentation fix.

Suggested change
@pytest.fixture(scope="session")
def remote_dataset():
"""Download and cache remote data once per test session."""
file_path = data_registry.fetch("sample_data.csv")
return pd.read_csv(file_path)
def test_remote_data_analysis(remote_dataset):
"""Test using remote dataset."""
result = analyze_dataset(remote_dataset)
assert result is not None
@pytest.fixture(scope="session")
def remote_dataset():
"""Download and cache remote data once per test session."""
file_path = data_registry.fetch("sample_data.csv")
return pd.read_csv(file_path)
def test_remote_data_analysis(remote_dataset):
"""Test using remote dataset."""
result = analyze_dataset(remote_dataset)
assert result is not None

@InessaPawson InessaPawson left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a valuable addition to the guide. I'd love to see it available as a community resource soon. I've left suggestions on what look like the blockers. Everything else could be refined in follow-up PRs.

Comment thread tests/package-data.md
Comment on lines +87 to +89
<!-- Is this true? -->
While these limits are not clearly documented,
most estimates are around 100 MB per file and 1 GB for the total project.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
<!-- Is this true? -->
While these limits are not clearly documented,
most estimates are around 100 MB per file and 1 GB for the total project.
The default limits are 100 MiB per file and 10 GB per project.
See PyPI's help page on [file size limits](https://pypi.org/help/#file-size-limit)
and [project size limits](https://pypi.org/help/#project-size-limit).

Comment thread tests/package-data.md
Comment on lines +282 to +284
```python
import pooch

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The snippet uses pd and xr without importing them. Perhaps, we should add them.

Suggested change
```python
import pooch
```python
import pandas as pd
import pooch
import xarray as xr

Comment thread tests/package-data.md
Comment on lines +55 to +59
<!-- This seems like a really BIG package 500mb?-->

A good rule of thumb is to have a handful of small files,
say no more than 10 files that are a maximum of 50 MB each.
Anything more than that, you will probably want to store online and download.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

500 MB is too big. A typical pure-Python wheel is usually under 1 MB. What about the following?

Suggested change
<!-- This seems like a really BIG package 500mb?-->
A good rule of thumb is to have a handful of small files,
say no more than 10 files that are a maximum of 50 MB each.
Anything more than that, you will probably want to store online and download.
A good rule of thumb is to keep bundled data to a handful of small files
totalling no more than about 10 MB. Anything larger than that, you will
probably want to store online and download on demand.

Comment thread tests/package-data.md

You want to avoid adding these files to your version control history (git) and draining the resources of PyPI.

<!-- I'm not sure what the intent of these example packages is - are they examples of packages that do this well? or poorly? -->

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1. Consider moving the list up under the pros.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

help wanted We welcome a contributor to work on this issue! thank you in advance! 🚀 ready-for-review

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

4 participants