Summary
ndi.cloud.batch_signed_url.BatchSignedUrlLookup._fetch_scope() at src/ndi/cloud/batch_signed_url.py:328 has the same silent-degradation failure mode as its MATLAB counterpart (VH-Lab/NDI-matlab#1010, filed alongside this one):
- The batch call may fail (network, auth, timeout). On failure the scope is marked as failed with a 60 s TTL (
DEFAULT_FAILURE_TTL_SECONDS), and every subsequent uid in that scope falls back to a per-member getFileDetails call.
- For a large file series (e.g. a lightsheet OME-Zarr level with 100k+ chunks) that is O(N) API calls where the fast path is O(1) per series. At ~100 ms per call, a 156k-member series takes 4+ hours.
- One transient TLS blip at start of run is enough to trigger the whole cascade. Home / residential networks reliably hit this.
What the code currently does
_fetch_scope (batch_signed_url.py:328):
try:
result = self._api_call(...)
except Exception as exc: # noqa: BLE001 - reported as a miss reason
# record failure with a 60s TTL, warn once, return empty
...
There is a "one retry per scope, ever" branch (batch_signed_url.py:263-290) but it only fires for the partial-map case — a scope that was populated but did not name the uid the caller asked about. On an initial fetch failure, no retry happens.
Suggested fix
Add exponential-backoff retry to the batch call inside _fetch_scope before treating it as a scope failure. Something like:
attempts = 3
delays = [1.0, 4.0, 16.0] # or read from a constant
last_exc = None
for i in range(attempts):
try:
result = self._api_call(...)
break
except Exception as exc:
last_exc = exc
if i < attempts - 1:
self._sleep(delays[i])
else:
# all attempts failed; record miss, warn once, return empty
...
A transient HTTP error then costs ~20 s of wall clock instead of hours. If all three attempts fail, the scope is legitimately unreachable and the fallback is correct.
The DEFAULT_FAILURE_TTL_SECONDS cache stays useful — it prevents hammering the endpoint again within the same sweep after a genuinely-failed scope.
Reproduction
- Cloud dataset with a file series of ≥100k members (e.g.
6ab034e430a0f8d0e461dfcc, a lightsheet OME-Zarr).
ndi.cloud.orchestration.downloadDataset(dataset_id, target, sync_files=False) from a network with occasional TLS instability.
- Either the batch call succeeds on first try (fast finish) or the initial call fails and the download runs for hours.
Impact on current work
The cloud integration tests in tests/test_pyramid_loader_integration.py pass in CI because GitHub's network is stable enough to keep the first batch call from failing. End-users on residential networks hit the fallback and see 4+ hour hangs on the same dataset.
Related
VH-Lab/NDI-matlab#1010 — parallel MATLAB fix (same retry, plus a second issue where batchSignedUrlLookup uses MATLAB webservices instead of curl).
Waltham-Data-Science/NDI-python#262 — original batch-signed-URL cache work.
Waltham-Data-Science/NDI-python#309 — partial-map retry (the existing retry branch that inspired this one).
Summary
ndi.cloud.batch_signed_url.BatchSignedUrlLookup._fetch_scope()atsrc/ndi/cloud/batch_signed_url.py:328has the same silent-degradation failure mode as its MATLAB counterpart (VH-Lab/NDI-matlab#1010, filed alongside this one):DEFAULT_FAILURE_TTL_SECONDS), and every subsequent uid in that scope falls back to a per-membergetFileDetailscall.What the code currently does
_fetch_scope(batch_signed_url.py:328):There is a "one retry per scope, ever" branch (batch_signed_url.py:263-290) but it only fires for the partial-map case — a scope that was populated but did not name the uid the caller asked about. On an initial fetch failure, no retry happens.
Suggested fix
Add exponential-backoff retry to the batch call inside
_fetch_scopebefore treating it as a scope failure. Something like:A transient HTTP error then costs ~20 s of wall clock instead of hours. If all three attempts fail, the scope is legitimately unreachable and the fallback is correct.
The
DEFAULT_FAILURE_TTL_SECONDScache stays useful — it prevents hammering the endpoint again within the same sweep after a genuinely-failed scope.Reproduction
6ab034e430a0f8d0e461dfcc, a lightsheet OME-Zarr).ndi.cloud.orchestration.downloadDataset(dataset_id, target, sync_files=False)from a network with occasional TLS instability.Impact on current work
The cloud integration tests in
tests/test_pyramid_loader_integration.pypass in CI because GitHub's network is stable enough to keep the first batch call from failing. End-users on residential networks hit the fallback and see 4+ hour hangs on the same dataset.Related
VH-Lab/NDI-matlab#1010— parallel MATLAB fix (same retry, plus a second issue wherebatchSignedUrlLookupuses MATLAB webservices instead of curl).Waltham-Data-Science/NDI-python#262— original batch-signed-URL cache work.Waltham-Data-Science/NDI-python#309— partial-map retry (the existing retry branch that inspired this one).