Skip to content

HateoasProvider doesn't support STAC best-practices and breaks with shared-folder item layouts #2404

Description

@kieranbartels

USGS recently restructured their stac repository and in doing so, temporarily broke their production version of pygeoapi. They used to store each Item in its own uniquely-named subdirectory (<id>/<id>.json), but restructured the repository to group Items under a shared items/ directory instead (items/<id>.json).

This change was driven by the catalog-layout best practices document, which states that a one-subdirectory-per-Item layout is only recommended when there are sidecar files alongside the Item. That layout should specifically be avoided when it would "regularly lead to a single Item in a directory," which was the case here, since these Items have no sidecar files.

The HateoasProvider.get_data_path builds rel: "item" links by splitting each source catalog.json/collection.json link's href on / and unpacking exactly three segments:

unused, path_ending, entry_type = link.split('/')

path_ending is then used as both the href and title for the item link. This assumes an Item's unique identifier is always the second segment. The HATEOAS provider worked correctly with the old structure, but with the new structure, path_ending resolves to the constant folder name items for every Item, and the real unique filename (the third segment) is discarded (shown below).

Image

Steps to Reproduce

  1. Serve a collection.json via the HATEOAS/STAC provider where Item hrefs follow <id>/<id>.json. Item links resolve correctly.
  2. Change the layout so Item hrefs follow items/<id>.json instead.
  3. Request the collection: GET /stac/stac-collection/<catalog>/<collection>?f=json.
  4. All rel: "item" links are now identical.

Or check out USGS's Pygeoapi development instance: https://labs-beta.waterdata.usgs.gov/api/gdp/pygeoapi/stac/stac-collection/nlcd/nlcd-FctImp

Expected behavior

Each rel: "item" link should have a unique href/title derived from the Item's actual filename, regardless of directory depth or whether Items share a parent directory.

Screenshots/Tracebacks

No exception raised. 40 Items all produce the same link:

Environment

  • OS: Windows
  • Python version: 3.12.13

Additional context

Suggested fix: derive the identifier from the basename of the link instead of a fixed-position segment.

for link in link_href_list:
  unused, path_ending, entry_type = link.split('/')
  newpath = os.path.join(baseurl, urlpath, path_ending).replace('\\', '/') # noqa
  
  if entry_type == 'catalog.json':
      child_links.append({
          'rel': 'child',
          'href': newpath,
          'type': 'application/json',
          'created': "-",
          'entry:type': 'Catalog'
      })
  elif entry_type == 'collection.json':
      child_links.append({
          'rel': 'child',
          'href': newpath,
          'type': 'application/json',
          'created': "-",
          'entry:type': 'Collection'
      })
  else:
      item_id = os.path.splitext(entry_type)[0]
      itempath = os.path.join(baseurl, urlpath, item_id).replace('\\', '/')  # noqa
      child_links.append({
          'rel': 'item',
          'href': itempath,
          'title': item_id,
          'created': "-",
          'entry:type': 'Item'
      })

Suggested fix for get_data_path's resolution cascade: add a fallback attempt for a flat sibling file, tried after the existing per-item-subdirectory attempt fails.

try:
    jsondata = _get_json_data(f'{data_path}/catalog.json')
    resource_type = 'Catalog'
except Exception:
    try:
        jsondata = _get_json_data(f'{data_path}/collection.json')
        resource_type = 'Collection'
        for key in ['license', 'extent', 'id']:
            if key in jsondata:
                content[key] = jsondata[key]
    except Exception:
        try:
            filename = os.path.basename(data_path)
            parent_dir = os.path.dirname(data_path)
            manifest = _get_json_data(f'{parent_dir}/collection.json')
            hrefs = {
                os.path.splitext(os.path.basename(link['href']))[0]: link['href']
                for link in manifest['links']
            }
            jsondata = _get_json_data(f'{parent_dir}/{hrefs[filename]}')
            resource_type = 'Assets'
        except Exception:
            msg = f'Resource does not exist: {data_path}'
            LOGGER.error(msg)
            raise ProviderNotFoundError(msg)

Backward-compatible with the existing <id>/<id>.json layout, and also correctly handles items/<id>.json grouping or any other nesting depth. Happy to provide before/after collection.json files and rendered output if useful.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions