USGS recently restructured their stac repository and in doing so, temporarily broke their production version of pygeoapi. They used to store each Item in its own uniquely-named subdirectory (<id>/<id>.json), but restructured the repository to group Items under a shared items/ directory instead (items/<id>.json).
This change was driven by the catalog-layout best practices document, which states that a one-subdirectory-per-Item layout is only recommended when there are sidecar files alongside the Item. That layout should specifically be avoided when it would "regularly lead to a single Item in a directory," which was the case here, since these Items have no sidecar files.
The HateoasProvider.get_data_path builds rel: "item" links by splitting each source catalog.json/collection.json link's href on / and unpacking exactly three segments:
unused, path_ending, entry_type = link.split('/')
path_ending is then used as both the href and title for the item link. This assumes an Item's unique identifier is always the second segment. The HATEOAS provider worked correctly with the old structure, but with the new structure, path_ending resolves to the constant folder name items for every Item, and the real unique filename (the third segment) is discarded (shown below).
Steps to Reproduce
- Serve a
collection.json via the HATEOAS/STAC provider where Item hrefs follow <id>/<id>.json. Item links resolve correctly.
- Change the layout so Item hrefs follow
items/<id>.json instead.
- Request the collection:
GET /stac/stac-collection/<catalog>/<collection>?f=json.
- All
rel: "item" links are now identical.
Or check out USGS's Pygeoapi development instance: https://labs-beta.waterdata.usgs.gov/api/gdp/pygeoapi/stac/stac-collection/nlcd/nlcd-FctImp
Expected behavior
Each rel: "item" link should have a unique href/title derived from the Item's actual filename, regardless of directory depth or whether Items share a parent directory.
Screenshots/Tracebacks
No exception raised. 40 Items all produce the same link:
Environment
- OS: Windows
- Python version: 3.12.13
Additional context
Suggested fix: derive the identifier from the basename of the link instead of a fixed-position segment.
for link in link_href_list:
unused, path_ending, entry_type = link.split('/')
newpath = os.path.join(baseurl, urlpath, path_ending).replace('\\', '/') # noqa
if entry_type == 'catalog.json':
child_links.append({
'rel': 'child',
'href': newpath,
'type': 'application/json',
'created': "-",
'entry:type': 'Catalog'
})
elif entry_type == 'collection.json':
child_links.append({
'rel': 'child',
'href': newpath,
'type': 'application/json',
'created': "-",
'entry:type': 'Collection'
})
else:
item_id = os.path.splitext(entry_type)[0]
itempath = os.path.join(baseurl, urlpath, item_id).replace('\\', '/') # noqa
child_links.append({
'rel': 'item',
'href': itempath,
'title': item_id,
'created': "-",
'entry:type': 'Item'
})
Suggested fix for get_data_path's resolution cascade: add a fallback attempt for a flat sibling file, tried after the existing per-item-subdirectory attempt fails.
try:
jsondata = _get_json_data(f'{data_path}/catalog.json')
resource_type = 'Catalog'
except Exception:
try:
jsondata = _get_json_data(f'{data_path}/collection.json')
resource_type = 'Collection'
for key in ['license', 'extent', 'id']:
if key in jsondata:
content[key] = jsondata[key]
except Exception:
try:
filename = os.path.basename(data_path)
parent_dir = os.path.dirname(data_path)
manifest = _get_json_data(f'{parent_dir}/collection.json')
hrefs = {
os.path.splitext(os.path.basename(link['href']))[0]: link['href']
for link in manifest['links']
}
jsondata = _get_json_data(f'{parent_dir}/{hrefs[filename]}')
resource_type = 'Assets'
except Exception:
msg = f'Resource does not exist: {data_path}'
LOGGER.error(msg)
raise ProviderNotFoundError(msg)
Backward-compatible with the existing <id>/<id>.json layout, and also correctly handles items/<id>.json grouping or any other nesting depth. Happy to provide before/after collection.json files and rendered output if useful.
USGS recently restructured their stac repository and in doing so, temporarily broke their production version of pygeoapi. They used to store each Item in its own uniquely-named subdirectory (
<id>/<id>.json), but restructured the repository to group Items under a shareditems/directory instead (items/<id>.json).This change was driven by the catalog-layout best practices document, which states that a one-subdirectory-per-Item layout is only recommended when there are sidecar files alongside the Item. That layout should specifically be avoided when it would "regularly lead to a single Item in a directory," which was the case here, since these Items have no sidecar files.
The HateoasProvider.get_data_path builds
rel: "item"links by splitting each sourcecatalog.json/collection.jsonlink'shrefon/and unpacking exactly three segments:path_endingis then used as both thehrefandtitlefor the item link. This assumes an Item's unique identifier is always the second segment. The HATEOAS provider worked correctly with the old structure, but with the new structure,path_endingresolves to the constant folder nameitemsfor every Item, and the real unique filename (the third segment) is discarded (shown below).Steps to Reproduce
collection.jsonvia the HATEOAS/STAC provider where Item hrefs follow<id>/<id>.json. Item links resolve correctly.items/<id>.jsoninstead.GET /stac/stac-collection/<catalog>/<collection>?f=json.rel: "item"links are now identical.Or check out USGS's Pygeoapi development instance: https://labs-beta.waterdata.usgs.gov/api/gdp/pygeoapi/stac/stac-collection/nlcd/nlcd-FctImp
Expected behavior
Each
rel: "item"link should have a uniquehref/titlederived from the Item's actual filename, regardless of directory depth or whether Items share a parent directory.Screenshots/Tracebacks
No exception raised. 40 Items all produce the same link:
Environment
Additional context
Suggested fix: derive the identifier from the basename of the link instead of a fixed-position segment.
Suggested fix for
get_data_path's resolution cascade: add a fallback attempt for a flat sibling file, tried after the existing per-item-subdirectory attempt fails.Backward-compatible with the existing
<id>/<id>.jsonlayout, and also correctly handlesitems/<id>.jsongrouping or any other nesting depth. Happy to provide before/aftercollection.jsonfiles and rendered output if useful.