Conversation
| * **[Open Science Framework (OSF)](https://osf.io)** - Comprehensive research platform | ||
| * **[Figshare](https://figshare.com)** - User-friendly with good visualization tools | ||
| * **[Dryad](https://datadryad.org)** - Focused on research data (subscription model for some features) | ||
|
|
There was a problem hiding this comment.
Would huggingface fit here or is this more for academic ? Perhaps https://docs.source.coop ?
There was a problem hiding this comment.
it's good you suggested these -- sorry for a big wordy response to a single suggestion, but I think it shows that we could say a little more about why we have what we have here.
I'll say what I think we might add, but first, re: your suggestions:
Huggingface is free, and convenient.
I have a couple of worries about it though. It will only exist (1) at most as long as the company exists, and that might not be long if the AI bubble pops, and (2) it might not even be that long if the company decides that the cost of maintaining all the data is not worth how many customers it gets them. I don't think this is just an academic thing, since a company could choose to be structured in such a way that they can't pull the rug on users whenever they want. Forgive me for being pedantic about it, I guess I'm sensitive as a failed academic 😇
Source Co-op looks cool! Maybe we should include them! Need to read about it more
re: the places we have listed now, I think there's a couple things we could do here:
- say a little bit more in the blurb about what our criteria are. E.g., is it free, is there some sort of guarantee on how long the data will be stored, is there a good UI / API / existing software tool that makes it easy to work with
- put in some sort of table for each of these that indicates how each place to store data stack up WRT to those criteria
Again, sorry for being wordy -- some of this stuff I say already in the actual test, but you're making me see how this could maybe be better organized
| * **[Google Cloud Platform](https://cloud.google.com/storage)** - Cloud Storage with strong AI/ML tool integration | ||
| * **[Linode](https://www.linode.com/products/object-storage/)** - Object Storage with straightforward pricing and developer-friendly tools | ||
|
|
||
| <!-- I don't understand how these platforms are different from things like figshare - can we clarify that? and how / why someone would pick these vs figshare / dryad? --> |
There was a problem hiding this comment.
This section is moving into data versioning which is important for scientific studies and reproducibility, but may be out of scope for a package. Having a small subset or testable data for a package makes a lot of sense.
| - **[Pachyderm](https://www.pachyderm.com/)** - Data pipeline platform with version control | ||
| ``` | ||
|
|
||
| <!-- I am not sure how this section relates to data stored in a package. I understand it's important, but does it belong in a page focused on how and where to store data for your Python package? It might be that I just don't understand as written!--> |
There was a problem hiding this comment.
I hear you -- I put this in thinking "I wish I knew about this earlier so we should mention it" but it is probably out of scope for the specifics of including data with your package
|
|
||
| ### Use Pytest fixtures for data access | ||
|
|
||
| Pytest fixtures provide a clean way to set up and share data across your test suite. They're especially useful for scientific packages where you need consistent access to test datasets. |
There was a problem hiding this comment.
Maybe this is just personal preference, but I think we should move pytest up and pooch down.
Pytest fits better with the whole 'testing your data thing' along with testing your code. Pooch seems like an additional tool or service to download data from an external source, I've not heard of it before.
There was a problem hiding this comment.
👍 We definitely should talk about pytest fixtures sooner if this page is about using data in tests.
@lwasser (at the risk of being even more annoying on this review 😅 ) looking at this again, I'm wondering if it would make sense to break this into two pages.
One on "data for tests" and another on "example data in your package", that would not live in the section on tests, but instead live in ... some section that doesn't exist yet.
I guess it would be some section that's "packaging", but more about, like library structure I guess? As oppossed to the nitty-gritty of publishing a package
I know a lot of times test data and example data overlap but it's striking me as odd to have content about example data in the tests/ section
ok I'll stop being noisy on the review 🤐
There was a problem hiding this comment.
One last thought: if I was writing a page that was only "data for tests", I'd write something like:
- as much as possible, use
pytest.fixturewith "fake data" you generate to test: this lets you avoid adding real data to your project, and helps you avoid writing code and tests that are tightly coupled, encouraging you to test the interface instead of implementation details. - for functionality where you absolutely need to test on real data, e.g. parsing specific file formats, include those in version control if possible, but avoid pushing the data to PyPI because it increases package size and strains the service
- in some cases you may include example data in your package (link to section on example data), and you can re-use this for tests, using fixtures, here's how...
- to test some functionality, e.g. fitting statistical models, you may not be able to avoid downloading relatively large amounts of data for your tests. In these cases you may find it convenient to publish your dataset to open data repositories and then download it as part of a set-up step for your project, that you should outline in your contributing.md. Tools exist for accessing these dataset [link to example data page again]
| <!-- I am not sure about this statement in terms of what it means and whether we have tools that consider standards or not we might also want to link to FAIR--> | ||
| ::: | ||
|
|
||
| ```{admonition} Field specific standards + metadata |
There was a problem hiding this comment.
A link to FAIR would be good. I haven't heard of the ones mentioned 😓
There was a problem hiding this comment.
+1 for saying that in general datasets should be FAIR.
I could add link to the neuro formats.
I know some of the astro ones too, e.g., FITS, and the loftily-named [Advanced Scientific Data Format](https://proceedings.scipy.org/articles/majora-212e5952-000. Wwe could ask astro people.
Would geo formats be worth a mention here too?
Or is all that info overload
|
YAY! ok there is feedback here - y'all i'm catchin up after travel and a workshop. I'll focus on incorporating everyone's comments and then we can do another round of review. |
d59e672 to
56c7011
Compare
|
|
||
| Here, you will learn about working with data for your scientific Python package. | ||
|
|
||
| ::{admonition} What you will learn |
There was a problem hiding this comment.
The "What you will learn" box is not rendering as a directive. Let's fix that.
| ::{admonition} What you will learn | |
| :::{admonition} What you will learn |
|
|
||
| * **movingpandas** is a pyOpenSci-accepted package that uses data to support some of its tutorials. [Here is an example of a tutorial that shows how to use MovingPandas to process bird migration data.](https://movingpandas.github.io/movingpandas-website/2-analysis-examples/bird-migration.html) | ||
|
|
||
| * [scikit-image:](https://github.com/scikit-image/scikit-image/tree/main/skimage/data) stores data within the package itself to be used for package examples. <i made this up> |
There was a problem hiding this comment.
| * [scikit-image:](https://github.com/scikit-image/scikit-image/tree/main/skimage/data) stores data within the package itself to be used for package examples. <i made this up> | |
| * [scikit-image](https://github.com/scikit-image/scikit-image/tree/main/skimage/data) |
We could expand in a follow-up PR.
| * core scientific-python packages | ||
| * scikit-learn: <https://github.com/scikit-learn/scikit-learn/tree/main/sklearn/datasets/data> | ||
|
|
||
| # Store Your Data in a Scientific Repository |
There was a problem hiding this comment.
H1 breaks the sidebar nav, and I'm not sure it's needed. Consider using H2 or H3.
Placing under "Where to store your data" with H3 seems reasonable.
| When creating data for your package, be aware of field-specific standards and formats. For example, in neuroscience, DANDI and NWB are common file formats used in the domain. So consider whether your package can / should support those formats if you expect users in the neuroscience space to use it. | ||
|
|
||
| ???Many pyOpenSci tools exist to address these standards or to provide interoperability because these standards don't exist ??? | ||
| see also: FAIR data |
There was a problem hiding this comment.
| see also: FAIR data | |
| Where possible, follow [FAIR data principles](https://www.go-fair.org/fair-principles/) so your example and test data are findable, accessible, interoperable, and reusable. |
|
|
||
| When creating data for your package, be aware of field-specific standards and formats. For example, in neuroscience, DANDI and NWB are common file formats used in the domain. So consider whether your package can / should support those formats if you expect users in the neuroscience space to use it. | ||
|
|
||
| ???Many pyOpenSci tools exist to address these standards or to provide interoperability because these standards don't exist ??? |
There was a problem hiding this comment.
What about:
| ???Many pyOpenSci tools exist to address these standards or to provide interoperability because these standards don't exist ??? | |
| Many pyOpenSci tools are built to support these field-specific standards, or to provide interoperability between formats where no common standard exists. |
| from importlib import resources | ||
| import my_package | ||
|
|
||
| with resources.open_text(my_package.data, 'example.csv') as f: |
There was a problem hiding this comment.
open_text is deprecated since Python 3.11. files() is the current API that works on 3.9+ and in the backport.
| with resources.open_text(my_package.data, 'example.csv') as f: | |
| with resources.files("my_package.data").joinpath("example.csv").open("r") as f: |
| def load_sample_data(): | ||
| file_path = data_registry.fetch("sample_data.csv") | ||
| return pd.read_csv(file_path) | ||
|
|
||
| def load_large_dataset(): | ||
| file_path = data_registry.fetch("large_dataset.nc") | ||
| return xr.open_dataset(file_path) |
There was a problem hiding this comment.
Fixes mixed indentation, will raise an error otherwise.
| def load_sample_data(): | |
| file_path = data_registry.fetch("sample_data.csv") | |
| return pd.read_csv(file_path) | |
| def load_large_dataset(): | |
| file_path = data_registry.fetch("large_dataset.nc") | |
| return xr.open_dataset(file_path) | |
| def load_sample_data(): | |
| file_path = data_registry.fetch("sample_data.csv") | |
| return pd.read_csv(file_path) | |
| def load_large_dataset(): | |
| file_path = data_registry.fetch("large_dataset.nc") | |
| return xr.open_dataset(file_path) |
| @pytest.fixture | ||
| def sample_data(): | ||
| """Load sample dataset for testing.""" | ||
| data_path = Path(__file__).parent / "data" / "sample.csv" | ||
| return pd.read_csv(data_path) | ||
|
|
||
| def test_data_processing(sample_data): | ||
| """Test uses the fixture automatically.""" | ||
| result = my_function(sample_data) | ||
| assert len(result) > 0 |
There was a problem hiding this comment.
Mixed indentation as above.
| @pytest.fixture | |
| def sample_data(): | |
| """Load sample dataset for testing.""" | |
| data_path = Path(__file__).parent / "data" / "sample.csv" | |
| return pd.read_csv(data_path) | |
| def test_data_processing(sample_data): | |
| """Test uses the fixture automatically.""" | |
| result = my_function(sample_data) | |
| assert len(result) > 0 | |
| @pytest.fixture | |
| def sample_data(): | |
| """Load sample dataset for testing.""" | |
| data_path = Path(__file__).parent / "data" / "sample.csv" | |
| return pd.read_csv(data_path) | |
| def test_data_processing(sample_data): | |
| """Test uses the fixture automatically.""" | |
| result = my_function(sample_data) | |
| assert len(result) > 0 |
| @pytest.fixture(scope="session") | ||
| def remote_dataset(): | ||
| """Download and cache remote data once per test session.""" | ||
| file_path = data_registry.fetch("sample_data.csv") | ||
| return pd.read_csv(file_path) | ||
|
|
||
| def test_remote_data_analysis(remote_dataset): | ||
| """Test using remote dataset.""" | ||
| result = analyze_dataset(remote_dataset) | ||
| assert result is not None |
There was a problem hiding this comment.
Same indentation fix.
| @pytest.fixture(scope="session") | |
| def remote_dataset(): | |
| """Download and cache remote data once per test session.""" | |
| file_path = data_registry.fetch("sample_data.csv") | |
| return pd.read_csv(file_path) | |
| def test_remote_data_analysis(remote_dataset): | |
| """Test using remote dataset.""" | |
| result = analyze_dataset(remote_dataset) | |
| assert result is not None | |
| @pytest.fixture(scope="session") | |
| def remote_dataset(): | |
| """Download and cache remote data once per test session.""" | |
| file_path = data_registry.fetch("sample_data.csv") | |
| return pd.read_csv(file_path) | |
| def test_remote_data_analysis(remote_dataset): | |
| """Test using remote dataset.""" | |
| result = analyze_dataset(remote_dataset) | |
| assert result is not None |
InessaPawson
left a comment
There was a problem hiding this comment.
This is a valuable addition to the guide. I'd love to see it available as a community resource soon. I've left suggestions on what look like the blockers. Everything else could be refined in follow-up PRs.
| <!-- Is this true? --> | ||
| While these limits are not clearly documented, | ||
| most estimates are around 100 MB per file and 1 GB for the total project. |
There was a problem hiding this comment.
| <!-- Is this true? --> | |
| While these limits are not clearly documented, | |
| most estimates are around 100 MB per file and 1 GB for the total project. | |
| The default limits are 100 MiB per file and 10 GB per project. | |
| See PyPI's help page on [file size limits](https://pypi.org/help/#file-size-limit) | |
| and [project size limits](https://pypi.org/help/#project-size-limit). |
| ```python | ||
| import pooch | ||
|
|
There was a problem hiding this comment.
The snippet uses pd and xr without importing them. Perhaps, we should add them.
| ```python | |
| import pooch | |
| ```python | |
| import pandas as pd | |
| import pooch | |
| import xarray as xr |
| <!-- This seems like a really BIG package 500mb?--> | ||
|
|
||
| A good rule of thumb is to have a handful of small files, | ||
| say no more than 10 files that are a maximum of 50 MB each. | ||
| Anything more than that, you will probably want to store online and download. |
There was a problem hiding this comment.
500 MB is too big. A typical pure-Python wheel is usually under 1 MB. What about the following?
| <!-- This seems like a really BIG package 500mb?--> | |
| A good rule of thumb is to have a handful of small files, | |
| say no more than 10 files that are a maximum of 50 MB each. | |
| Anything more than that, you will probably want to store online and download. | |
| A good rule of thumb is to keep bundled data to a handful of small files | |
| totalling no more than about 10 MB. Anything larger than that, you will | |
| probably want to store online and download on demand. |
|
|
||
| You want to avoid adding these files to your version control history (git) and draining the resources of PyPI. | ||
|
|
||
| <!-- I'm not sure what the intent of these example packages is - are they examples of packages that do this well? or poorly? --> |
There was a problem hiding this comment.
+1. Consider moving the list up under the pros.
This pr is a rework of #110 . Let's plan to run a sprint on this pr in the next few weeks to see if we can get to a place where it's good.
NOTE: I added some code examples that I literally found online and DID NOT TEST. so we will want to definitely test what is there before merging.
ALSO - because this topic is not my expertise, in some cases I ran with a section after a Google search and tried to flesh it out, but it also could be off. Any feedback is welcome on this!!