Skip to content

batch_signed_url: initial _fetch_scope failure cascades to O(N) per-member API calls; add retry-with-backoff before falling back #322

Description

@stevevanhooser

Summary

ndi.cloud.batch_signed_url.BatchSignedUrlLookup._fetch_scope() at src/ndi/cloud/batch_signed_url.py:328 has the same silent-degradation failure mode as its MATLAB counterpart (VH-Lab/NDI-matlab#1010, filed alongside this one):

  • The batch call may fail (network, auth, timeout). On failure the scope is marked as failed with a 60 s TTL (DEFAULT_FAILURE_TTL_SECONDS), and every subsequent uid in that scope falls back to a per-member getFileDetails call.
  • For a large file series (e.g. a lightsheet OME-Zarr level with 100k+ chunks) that is O(N) API calls where the fast path is O(1) per series. At ~100 ms per call, a 156k-member series takes 4+ hours.
  • One transient TLS blip at start of run is enough to trigger the whole cascade. Home / residential networks reliably hit this.

What the code currently does

_fetch_scope (batch_signed_url.py:328):

try:
    result = self._api_call(...)
except Exception as exc:  # noqa: BLE001 - reported as a miss reason
    # record failure with a 60s TTL, warn once, return empty
    ...

There is a "one retry per scope, ever" branch (batch_signed_url.py:263-290) but it only fires for the partial-map case — a scope that was populated but did not name the uid the caller asked about. On an initial fetch failure, no retry happens.

Suggested fix

Add exponential-backoff retry to the batch call inside _fetch_scope before treating it as a scope failure. Something like:

attempts = 3
delays = [1.0, 4.0, 16.0]  # or read from a constant
last_exc = None
for i in range(attempts):
    try:
        result = self._api_call(...)
        break
    except Exception as exc:
        last_exc = exc
        if i < attempts - 1:
            self._sleep(delays[i])
else:
    # all attempts failed; record miss, warn once, return empty
    ...

A transient HTTP error then costs ~20 s of wall clock instead of hours. If all three attempts fail, the scope is legitimately unreachable and the fallback is correct.

The DEFAULT_FAILURE_TTL_SECONDS cache stays useful — it prevents hammering the endpoint again within the same sweep after a genuinely-failed scope.

Reproduction

  • Cloud dataset with a file series of ≥100k members (e.g. 6ab034e430a0f8d0e461dfcc, a lightsheet OME-Zarr).
  • ndi.cloud.orchestration.downloadDataset(dataset_id, target, sync_files=False) from a network with occasional TLS instability.
  • Either the batch call succeeds on first try (fast finish) or the initial call fails and the download runs for hours.

Impact on current work

The cloud integration tests in tests/test_pyramid_loader_integration.py pass in CI because GitHub's network is stable enough to keep the first batch call from failing. End-users on residential networks hit the fallback and see 4+ hour hangs on the same dataset.

Related

  • VH-Lab/NDI-matlab#1010 — parallel MATLAB fix (same retry, plus a second issue where batchSignedUrlLookup uses MATLAB webservices instead of curl).
  • Waltham-Data-Science/NDI-python#262 — original batch-signed-URL cache work.
  • Waltham-Data-Science/NDI-python#309 — partial-map retry (the existing retry branch that inspired this one).

Activity

  1. stevevanhooser commented on Sep 28, 2026

    @stevevanhooser
    ContributorAuthor

    Closing — the ask here (retry-with-backoff inside _fetch_scope before falling back per-member) landed on claude/lightsheet-zarr-ndi-viewer-djp5vk in commit 87d0d31 (batch_signed_url: retry the batch call before falling back per-member). Same shape as the fix sketch in the description: exponential-backoff attempts before treating the scope as failed, and DEFAULT_FAILURE_TTL_SECONDS still guards the fallback within a sweep.

    Follow-on landed alongside: a persistent (on-disk) signed-URL cache so a reopen of the same dataset pays the ~85 min async signed-URL-set-job cost once rather than every viewer open. NDI-python commit 18ec747 (src/ndi/cloud/signed_url_disk_cache.py, wired through BatchSignedUrlLookup(disk_cache=True) and forgotten on S3 403 in fetch_cloud_file). MATLAB mirror at VH-Lab/NDI-matlab@27661a3. Both implementations write byte-for-byte identical files, so a cache dir written by MATLAB reads clean from Python and vice versa.

    Both are on PR #320.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions