Repository navigation
Gene pyramid viewer: panels, controls, cloud credentials, and one-based tile files - #291
Merged
Merged
Conversation
Three things the viewer could not do. Two of them looked like flags the caller had forgotten to pass, which is its own kind of bug report. CELL OUTLINES. The cells document has carried contours since it was first written, readContourFile has parsed them since it was ported, and nothing drew them: openPyramid added centroids as Points and stopped. readContours is the missing document layer over readContourFile -- it finds contours.bin in the document and, crucially, applies contour_reference. Vertices are usually stored relative to their cell's centroid, because that is what makes them fit in int16, so a caller that drew readContourFile's output directly would put every outline in a small cluster near the origin: a picture of nothing, produced without an error. The tests assert the two references separately AND assert that they differ, since a reader that ignored the field entirely would pass a test that only checked one. An empty polygon comes back as (0, 2) rather than being dropped, so row i here stays row i of cells.tsv; dropping it would slide every later outline onto the wrong cell. napari will not take a zero-vertex shape, so those are filtered at the layer instead. GENES AND DENSITY. Both were fixed at launch, so changing either meant quitting and reopening the session. Neither needs that: a level is summed over the selected genes into a 2D array, so the shape does not change when the selection does and only the layer's data has to be swapped. The panel lives in its own module because it needs magicgui, which napari brings and a caller composing layers by hand may not; a missing magicgui now costs the panel rather than the picture. A gene name that matches nothing refuses the whole apply rather than drawing the ones that did match, because a partial picture looks exactly like a complete one. --name. The image layer was named from the pyramid's label, which the ingest usually took from the filename, so it announced the section rather than what was being shown -- actively misleading once --genes narrows it. readContours was very nearly a second copy of the binary parser: the grep that said no reader existed was for lowercase "contour", which does not match readContourFile. It delegates. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Rebuilt the controls to match the widgets the FICTURE viewer in the bscholl project already uses, because those are the shapes people here know and a different idiom for the same job is a cost with no return. The magicgui text box asked you to already know a gene name; the point of a picker is not having to. ADD GENE. Type-ahead QComboBox with a MatchContains completer over the whole list, showing "SYMBOL (ACCESSION)" so an unnamed or duplicated symbol is still selectable -- this opossum annotation repeats 324 symbols in its first 2,000. Adding creates its own additive layer rather than replacing the base, which is what makes two genes comparable; adding the same gene twice replaces its own layer instead of stacking. Every row a symbol matches is summed, since taking the first would show part of a gene's signal and look complete. DISPLAY. Radio buttons, not a checkbox: two exclusive states, so each label can say what it IS instead of naming one and leaving the other implied. Switching swaps only the layer's data, so the camera, the zoom and the level survive. It rebuilds with the SAME gene rows the base layer was drawn from, remembered on the layer -- recomputing with all genes would quietly widen what is shown. CELL TYPES. The cellTypeLabels documents were being written and never read: there was a writer and no reader at all. readCellTypeLabels is the reader, and it places labels BY cell_index rather than by file order, because cell_index is written explicitly for exactly that reason and a reordered file would otherwise give every cell its neighbour's type. An unlabelled cell stays as an empty string rather than being dropped, which would shift every later row. Each labeling gets its own checkbox group rather than being merged: a cellbin carries a transferred atlas call AND clusterings, and they are not interchangeable. Hiding sets ALPHA rather than deleting points, because the points layer's row order is the contract with cells.tsv and with every other labels.tsv beside it; deleting would break the next labeling the user toggles. Two labelings therefore filter jointly instead of the last click winning. The viewer now takes cells_doc, which it needs because the overlay arrives as plain arrays with no way back to the document the labels depend on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The gene picker listed one entry per pyramid COLUMN, and a real annotation names the same gene on several: this opossum list repeats 324 symbols in its first 2,000, because a symbol can carry several accessions. So the same gene appeared many times over and the reader was left to guess which one was the gene. geneIndex now collapses the list to one entry per symbol and keeps every row it names; the layer sums them, since taking the first would draw part of a gene's signal and look like all of it. The combobox becomes a filterable list with a checkbox each, which is what the FICTURE viewer next door already does. Ticking adds that gene's own additive layer in its own colour -- one map for everything made two genes into a single picture that neither of them was -- and unticking removes it; closing the layer in napari's own list unticks the box, so the two cannot disagree. Two labelings that group the cells identically and differ only in what the groups are called are now named once and drawn once. That is not a hypothetical: a subclass call transferred onto a clustering is a RENAMING of those clusters, so it agrees with the clustering by construction, and showing both invites the reader to treat one as corroborating the other. labelPartition compares MEMBERSHIPS, since comparing names would miss exactly this case, and supervised calls are considered first so the labeling carrying biological names is the one that stays. Panels now read down the right-hand side in the order they are used: display mode, cell types, then the long scrolling gene list. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
Comparing partitions for EQUALITY was too strict for the case it was written for. A subclass call transferred cell by cell onto a clustering agrees with it almost everywhere and disagrees on a handful of boundary cells: not a second opinion, the same opinion with noise -- but not the same partition either, so both were drawn and the reader was back where they started. labelAgreement measures it instead. For each group of one labeling, the majority value of the other is what it predicts, and the agreement is the fraction that prediction is right for; a per-cluster renaming scores 1.0. It is DIRECTIONAL, since a fine clustering determines a coarse call without the reverse holding, so both directions are tested and the larger is reported. The number is quoted in the panel every time rather than applied silently, so a pair that only nearly agrees reads as a pair that only nearly agrees. selectLabelings is pure and holds the decision, which is the part that can be wrong. --labels overrides it outright: a comma list of names, or 'none'. A name that is not there comes back as missing and is reported, because a typo that silently showed everything would look like the flag not working. labelPartition goes: agreement of 1.0 in both directions is partition equality, and two ways to ask the same question is one too many. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
VH-Lab/NDI-matlab#977 renames ndi.gui.app.GeneIngest to ndi.gui.app.GEFManager, so the matlab_path recorded here points at a file that no longer exists. test_every_recorded_matlab_path_points_at_a_real_file catches exactly that, and it caught this. The decision_log gains the half that is no longer symmetrical in the usual direction: the app's new View button draws nothing in MATLAB, it runs napariViewGEF from this package. So what stays not_yet_ported is the ingest and the listing, and the viewing the MATLAB app offers is already this side's -- reached across a process boundary rather than reimplemented. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
The wait before the window opens was not NDI and was not the pyramid.
The image ladder is lazy by construction -- layerSpec resolves tile paths
and reads no tile bytes -- so every second of it was the CELLS, which are
read whole. Measured on the opossum section's shape (493,126 cells, 24
vertices each, a 26 MB cells.tsv and a 49 MB contours.bin):
3.4s cells.tsv parsed
3.4s cells.tsv parsed AGAIN, inside readContours
2.0s one np.stack per cell to split contours.bin
3.3s one np.column_stack per cell to add the centroids back
13.8s list(zip(row, col)) per polygon in openPyramid
0.4s list(zip()) for the centroids
Nothing there is doing real work per cell; it is half a million Python
round trips through numpy for operations that are one array each.
readContourFile now builds one (total, 2) array and hands back views into
it, which is also what lets readContours place every vertex in a single
addition: np.repeat gives each vertex its own cell's centroid, so a
per-cell offset becomes a whole-array operation. readContours takes the
centroids the caller already parsed instead of re-reading cells.tsv. The
viewer uses np.column_stack per polygon and one for the whole centroid
table. Same numbers out, and the placement is now pinned by tests --
including the empty-file case, where np.split would otherwise hand back
one polygon where there are none.
2.5s readCells
0.7s readContours
1.6s world transform
0.0s centroids
WHAT IS LEFT IS NAPARI, and this cannot make it fast: add_shapes
triangulates every polygon as it takes them, and there are half a
million. So the launch now narrates itself to stderr, a line per stage
with its elapsed time, and names that step as the long one before it
starts. A program that makes you wait owes you an account of the wait.
NDI_GENEPYRAMID_QUIET=1 turns it off.
Also: the density-versus-counts explanation is hover text on the two
radio buttons now, not a paragraph standing under them. It is read once.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
The list said which genes exist and nothing about which ones carry the section. Each row now shows the reads that gene contributes, summed over the same variant rows the layer sums, so the number in the list is the number the picture is drawn from. readGeneTotals reads gene_totals.tsv off the pyramid, which had a writer and no reader. Older pyramids do not carry that file: it comes back as None rather than as zeros, because a gene with no reads and a gene whose reads were never recorded are different things, and the column and the band are dropped rather than invented. A totals file whose length disagrees with the gene list is refused for the same reason -- indexing it by gene row would put one gene's reads on another. The two-ended band is ported from the FICTURE viewer's abundance panel, including the parts that make it correct rather than merely plausible. It is a percentile of the READS, not of the list, and on a real section those are wildly different: expression is heavy-tailed, so the top half of the reads is a percent or so of the genes. An entry occupies an interval of the cumulative sweep and sits at its MIDPOINT, since either endpoint would let one dominant gene straddle a cut on the strength of its own bulk; tied totals share one midpoint computed over the whole group, or ties break on sort order and two genes with identical counts land on opposite sides of the same threshold. That arithmetic has a consequence which reads as a bug the first time: a gene carrying 30% of all reads has the interval [70, 100], so a 1% top cut cannot reach it -- correctly, since dropping it would remove 30% of the reads. The readout says so, and names the cut that would work, rather than letting the control look inert. It also names what a top cut did drop, because "you just dropped MBP" is the useful feedback and "1,204 genes hidden" is not. The band filters the LIST; it composes with the text box rather than replacing it, since narrowing to the loud genes and then searching within them is how both get used. superqt's range slider when it is there, two spin boxes when it is not, and the same presets as before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
The abundance band leaves you wanting this. Narrow to the top 10% of reads and you get a few hundred genes in alphabetical order, which hides the one thing you narrowed down to see -- which of them are loud. countsOrder is stable, and that is its whole content: the list arrives alphabetical, so a stable sort leaves equal counts in alphabetical order instead of an arbitrary one, and at the bottom of a real section thousands of genes tie. It is pure and tested for exactly that, because a later edit dropping kind="stable" would look harmless and read as a list that reshuffles itself. The list is REBUILT rather than reordered. Moving thirty thousand items one at a time is quadratic and sorting them through a Python comparison is no cheaper than making them again, while a rebuild is the same single batched pass the panel already pays once. Two things follow from that and both are now explicit: the abundance mask is computed once in the original order, so each item carries its original position and the filter reads that rather than its row in the widget; and the ticks come back from the LAYERS rather than from remembered state, since a drawn layer is what a tick means and reading it is what stops the boxes and the picture from disagreeing across a sort. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
THE BAND IS ONE ROW NOW: a label, the slider, nothing else. The header, the preset buttons and the readout paragraph were four lines of chrome above the thing the panel is for. What the band is, and what it just did -- how many genes it kept, what share of the reads, which loud genes a top cut dropped, and the case where one gene is too loud for a small cut to reach -- is hover text on the slider. The explanation is read once; the list is read every time. The one number worth keeping in sight goes into the header line that was already there, since "1,204 of 30,434 genes" is the same fact that line already states, narrowed. SORTING IS A HEADER ROW: Gene and Counts as flat buttons that read like a table's header and behave like one. Click to sort by that column, click again to reverse. Four orders, which is what a ranked list needs; a menu would have listed all four and made the reader pick wording rather than direction. countsOrder gains a direction, and ascending is a separate sort rather than the descending one reversed -- reversing would reverse the ties with it, so a tied group would read alphabetically one way and backwards the other. At the bottom of a real section thousands of genes tie, so that is most of the list. Asserted in both directions, including that ascending is NOT the reverse of descending. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
Dropping a redundant labeling automatically was wrong in both directions. Below the agreement threshold it did nothing and the reader still saw two of them -- which is what happened here. Above it, the labeling vanished with no way to bring it back. So nothing is dropped now: every labeling the cells document has is shown, the pair that says the same thing is NAMED with the number, and the reader unticks one. That is the same information with the decision left where it belongs. Each labeling is ONE ROW: its own checkbox, the class count, and a toggle that opens its classes. A cellbin carries a couple of labelings with dozens of classes between them, and listing every class of every one filled the dock before anyone had asked to see them. The unlabelled count and the agreement number move to the row's hover text. Switching a labeling off DROPS ITS CONTRIBUTION TO THE FILTER rather than hiding the cells it named. That is the whole point of the switch: an unwanted second opinion should stop having an opinion, not start hiding everything it had named. The joint filter still holds across the labelings that are still on, so a cell is dimmed when any switched-on labeling has its class unticked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
By the time addAllPanels runs the viewer is on screen and the image is drawn. A panel that raised unwound out of openPyramid before napari.run(), so the window appeared and the process exited -- which is seen as napari opening briefly and closing, with the reason lost unless somebody was watching stderr. The docstring already stated the principle for a missing Qt and then applied it only to that import. Each panel now builds independently and a failure costs that panel, named, with its traceback. This is not hypothetical: every panel reads files out of the database to build itself -- genes.tsv and gene_totals.tsv for the genes, one labels.tsv per labeling for the cell types -- and those reads fail for reasons that have nothing to do with the image, a cloud-backed session among them. None of them is a reason to refuse to draw the section. AND --labels NEVER WORKED. addAllPanels called addCellTypePanel without forwarding labelings, in every commit since the flag was added. The flag was parsed, resolved, threaded through openPyramid and then dropped at the last hop, so choosing labelings in the CLI or in the GEF Manager's dialog silently did nothing and the panel always took its default. It went unnoticed because every other layer mentions the argument, so grepping for it finds it everywhere it is not being used. Both that forwarding and the panel isolation are now pinned by tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
ndi.cloud.profile documents the AES-128/CBC file backend as its DEFAULT and nothing installed it. cryptography was in neither the dependencies nor any extra, so on a clean install _detect_backend fell through to the in-memory backend -- which persists nothing and reads nothing -- and said so only through a logger warning. A saved cloud password could therefore never be loaded, which is not a degraded mode so much as the feature being absent. Not an extra. ndi.cloud is part of the package rather than an opt-in surface like gui or napari, and a default backend that nothing installs is not a default. keyring stays optional: it is preferred when present, but it is the OS-keychain path, and cryptography is the portable one that MATLAB's AES file interoperates with. Worth being clear about what this does NOT fix: the profiles list itself is plain JSON and _load_from_disk reads it regardless of backend, so an empty list_profiles() is a question about where ~/.ndi is, not about this. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
jsonencode of a 1x1 struct produces an object; of a 1xN struct array, an array. So NDI-matlab's Profiles field has a different SHAPE depending on how many profiles exist, and the two languages share this file by design. MATLAB's reader has normalizeProfiles for exactly that. Python had no counterpart: it iterated the object as a list, got the dict's KEYS, and skipped every one as "not a dict". The result was zero profiles and no error -- the file was found, parsed, and silently read as empty -- for the single-profile case, which is what every new user has after making their first. That is the worst possible case to get wrong, and it is the one that was. _as_profile_list normalises both shapes and refuses the rest. Tests cover one, several, absent, null and nonsense, and were checked to fail against the old reader before the fix rather than merely passing after. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
ndi.cloud.profile is the documented place to keep an NDI Cloud password,
but nothing in the authentication path ever read it. CloudConfig.from_env
and authenticate() were environment-only, so every implicit login went
token -> NDI_CLOUD_USERNAME/PASSWORD -> raise, with no step in between.
The visible symptom: opening a downloaded cloud dataset in the pyramid
viewer dies inside the ndic:// file handler with
FileAccessError: ... No valid token and no credentials available.
Set NDI_CLOUD_TOKEN or NDI_CLOUD_USERNAME/NDI_CLOUD_PASSWORD.
while a working default profile sat on disk. Every @_auto_client API call
had the same hole.
MATLAB's authenticate.m has had the equivalent step all along --
authenticatedWithSecret, reading the MATLAB Vault -- so this is a parity
gap rather than a new feature. It goes after the environment step, not
before: an explicitly exported NDI_CLOUD_USERNAME is a deliberate
override and should keep winning over what is on disk.
The profile's Stage picks the API URL, but only when neither
NDI_CLOUD_URL nor CLOUD_API_ENVIRONMENT is set -- either of those means
the caller has already chosen. _profile_credentials() swallows its own
failures (no profile file, no default, a locked keyring) so a missing
credential surfaces as the caller's own error rather than a traceback
from an unrelated call site.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
NDI-matlab runs every secret key through matlab.lang.makeValidName before
using it as a struct field. It has to: reading the secrets file back
through jsondecode requires each JSON key to be a valid MATLAB
identifier. Python has no such constraint, so _safe_field substituted
underscores for spaces -- the reasonable guess, and the wrong one.
makeValidName DELETES whitespace and capitalises the following letter.
So MATLAB wrote
NDICloud41269713c8f30937_40d3dfe383e33674
and Python looked for
NDI_Cloud_41269713c8f30937_40d3dfe383e33674
Neither implementation could read the other's secrets, in either
direction, which is the entire point of a shared ~/.ndi store. The
symptom was a KeyError from get_password on a profile that had just
listed correctly, with its password sitting in the file two lines away.
_make_valid_name is the port: whitespace deleted and the next letter
capitalised, remaining non-alphanumerics to underscores, a leading 'x'
when the result would not start with a letter, keywords suffixed,
truncation at namelengthmax. The keyword and length rules are
unreachable for NDI's own keys; a port that implements four of five
rules is a trap for whoever needs the fifth.
Reads fall back to the old name, so a password saved by an older
NDI-python still opens; the next write moves it across and drops the
stale copy, and remove takes both. Nothing writes the old name again.
Note the capitalisation applies to UIDs too: "NDI Cloud f68d..." becomes
NDICloudF68d..., and half of Python's uuid4 UIDs begin with a-f. MATLAB
does the same thing, so tidying it up would break the interop rather
than improve it. There is a test saying so.
The last test builds a MATLAB-shaped store from the outside -- a single
profile as a JSON object, a <16hex>_<16hex> UID, a MATLAB field name --
and reads the password back through the public API, which covers the
whole handoff rather than this one field.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
A profile carries a Stage, prod or dev, and it decides which API the
requests go to. NDI-matlab settles this in profile.switchProfile --
setenv('CLOUD_API_ENVIRONMENT', prof.Stage)
-- which api/url.m then reads for every call. Python applied the stage
only on the credential path in authenticate(), and that is not the same
thing.
login() exports NDI_CLOUD_TOKEN. So the SECOND CloudConfig.from_env() in
a process carries a valid token, authenticate() returns at its first step
and never reaches the credential path, and the config keeps the prod
default URL it was built with. A dev profile therefore logged in against
dev and issued everything afterwards against prod. Each on-demand pyramid
tile builds its own client, so a viewer session failed from the second
request onwards with
Not found (HTTP 404): {"error":"Dataset ... does not exist or user
does not have access"}
which reads as a permissions problem rather than a wrong-server one --
the two are indistinguishable from the response, since a dataset that is
not on this server is exactly as absent as one you cannot see.
The stage now resolves in CloudConfig.from_env, which is the factory for
"build from ambient state" and so the right place for it. Precedence:
NDI_CLOUD_URL, then CLOUD_API_ENVIRONMENT, then the profile's Stage, then
prod -- so both environment variables still override, and an explicitly
constructed CloudConfig is untouched by construction. An unrecognised
stage falls back to prod rather than building a URL nothing serves.
authenticate() keeps only the credential half. Nineteen tests, including
the short-circuit itself: a config built with a token already exported
still points at dev.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
did.document/addFileSeries takes "ONE-BASED member numbers" and refuses anything else -- "Member indices must be positive integers (one-based)". The pyramid writer named its first tile tile.bin_0, putting every pyramid one step outside the convention its own storage layer assumes. That is not cosmetic. NDI-matlab's uploader walks a legacy NAME# series as NAME1, NAME2, ... and treats the first absent name as the end of the series, because that gap was the only termination signal it had. A zero-based pyramid hands it a series beginning at a name it never probes. The tile INDEX stays zero-based: it is a grid position, and index_order defines it as row*tile_columns + column. Only the file SUFFIX moves, so the file holding tile N is tile.bin_<N + tile_index_origin>. The origin is RECORDED on the document, not inferred. Pyramids already built name their first tile tile.bin_0 and carry no value, which reads as 0 and keeps them opening -- which matters, since the sessions being demonstrated were built before this. Inferring it from the stored names is not possible: a sparse pyramid whose first tile is empty has no _0 under either convention, and a wrong guess shifts every tile one cell along its row. An image that is quietly wrong is worse than one that is missing. Both directions are covered. tileFileName maps a grid index to a name; tileIndexFromName maps back, for the two readers that walk the stored names and work out (row, column) from the suffix -- doc_gene.exportRegion and doc_gene_export.readAsRecords. Reading the suffix AS the index there would have moved every tile one cell right, and the last tile of a row into the next row. The round-trip test builds a pyramid whose 2x2 grid stores tiles 1 and 4 with nothing between them, then walks it the way the uploader does and asserts the walk stops at the hole after one file. That sparseness is the normal case, not an exotic one: neither writer emits a tile with no data. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
The completeness test flagged four MATLAB functions with no entry. Three are this branch's tile-naming helpers; the fourth, hdf5Contents, has been unrecorded since it was added, and went unnoticed because NDI-python's CI runs on main pushes and pull requests and this branch has had neither. tileFileName and tileIndexFromName are ported and recorded as ordinary entries. The indexing note is the part worth having written down: the tile INDEX is zero-based on both sides, because it is a grid position and index_order defines it as row*tile_columns + column, while the file SUFFIX is one-based because a DID file series names its first member NAME_1. Those are two different numbers and the contract now says so. tileIndexOrigin is marked not_applicable. It exists in MATLAB only because reading an optional numeric field there takes isfield, isnumeric and isscalar guards, where Python has one .get with a default; a Python function wrapping that would be a wrapper around nothing. Its BEHAVIOUR is mirrored and tested on both sides -- absent, empty, non-numeric or non-scalar all read as 0 -- which is what the contract is for. hdf5Contents is marked not_yet_ported with a reason rather than back-filled: it answers a question asked interactively, at a MATLAB prompt, before committing to an ingest that NDI-matlab drives. Nothing in the Python package calls it. h5py would make the port easy; it is the verdict wording that would need mirroring, not the walk. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
…cs-hdf5-format-9jn3l9-viewer # Conflicts: # src/ndi/cloud/profile.py # src/ndi/fun/ndi_matlab_python_bridge.yaml # src/ndi/gui/app/genepyramid/multiscale.py # src/ndi/gui/ndi_matlab_python_bridge.yaml # src/ndi/util/ndi_matlab_python_bridge.yaml
main retired both spellings these entries used, with the replacement named in the failure text: not_applicable -> matlab_only / ported_differently / retired not_yet_ported -> porting_deferred / ported_differently tileIndexOrigin is matlab_only. The documented line is whether the replacement lives in this repo or the language provides it: reading an optional numeric field with a default takes isfield, isnumeric and isscalar guards in MATLAB and so earns a name, but in Python it is one dict.get from the standard library, written inline in tileFileName and tileIndexFromName. That is the "use tempfile" case, not the "use our SyncIndex class" one. hdf5Contents is porting_deferred: a counterpart is wanted, nobody has written it. The section banner carried the retired spelling too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
git hash-object returns the hash of a file's CONTENTS. The field wants a
COMMIT -- the point in NDI-matlab's history the entry was examined
against -- because the question it exists to answer is
git log <matlab_last_sync_hash>..HEAD -- <matlab_path>
and a blob is not a point in history to walk from. This is the mistake
tests/test_matlab_bridge_hashes.py was written for after 35 entries
carried blob hashes; the check could not run here, because the local
NDI-matlab checkout is shallow and it skips rather than reporting every
commit as missing.
The five entries now name each file's last-touching commit: 10687267a
for the three tile helpers, 742ac0d81 for hdf5Contents, and a328f306e
for GEFManager, whose recorded 8edda6d was a commit but a stale one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
The bridge CI job pairs the two repos by branch name, so this is the first run that has actually compared these YAMLs against the NDI-matlab branch they describe. Nine entries drift, all of them because THIS PR pair moved the MATLAB file. Two are real ports, not bumps: ndi.cloud.profile.reset now restores the detected backend. Python had the same hole NDI-matlab did: use_backend is a test hook, the override it sets is singleton state that lives for the whole process, and reset cleared everything except that. One test selecting the memory backend left every later set_password writing to a dict discarded at exit -- no error, and get_password returned what it had just stored, so a password looked saved when it was not. That is not a hypothetical; it cost an hour of debugging on the MATLAB side tonight before the cause was found. ndi.preferences gains GUI/GEFManager/ViewerLauncher, mirroring NDI-matlab 362060795, so the shared preference store round-trips a value MATLAB writes. The rest are examined-and-bumped, and the guard is right to demand a reason for each: fromCellBin and fromGEF changed help text only (the GeneIngest -> GEFManager rename, and recordSource now naming fileReference), so their decision_logs say so and name b589db9bb. makeSourceFile's says what actually changed -- describing a file the database does not hold is now a fileReference, a class with no files section, and only attachFile=true still produces a generic_file. The old shape was a generic_file with nothing attached, which is invalid; nothing enforced it until DID-matlab #182 closed a fail-open. A Python port must follow that, so the entry says it. Also fixed from the same run: the hashes for the three tile helpers, hdf5Contents and GEFManager were git hash-object output, which is a BLOB. The field names a commit, or `git log <hash>..HEAD -- <path>` has nothing to walk from. Full suite 32 failed / 5229 passed, against 40 / 5106 for main in the same working tree. The failure-list diff has eight this branch fixes and none it adds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
test (3.10/3.11/3.12) failed on two of this branch's own tests, and both
passed here. The difference is the install: addAllPanels refuses to build
anything when qtpy.QtWidgets will not import, and CI's ndi_install.py
--dev brings no Qt binding, while this container happens to have PyQt5.
So on CI the function returned {} before touching a panel, and the two
tests got an empty dict -- "the cell type panel was never built" and
KeyError: 'display'.
The tests were wrong, not the guard. What they check is the
ORCHESTRATION around the panels: that --labels reaches the cell type
panel, and that one panel throwing neither stops the ones after it nor
unwinds into openPyramid after the viewer is on screen. None of that is
about Qt.
The probe is now _importQtWidgets rather than an inline import, so a test
can stand in for it. Worth noting what the old shape cost: those two
tests asserted their behaviour anywhere a binding happened to be
installed and silently asserted nothing everywhere else -- passing either
way, which is the failure mode a test is supposed to prevent.
The guard's own path was marked no-cover because an inline import cannot
be reached from a test. It can now, so it has one: no binding builds
nothing, attempts nothing, and prints a line naming what still works.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
The bridge job's ndi_common sync check runs with NDI_COMMON_SYNC_STRICT=1
against the paired NDI-matlab checkout, and found three files this PR
pair added on the MATLAB side only:
database_documents/data/fileReference.json
database_documents/data/fileReference.md
schema_documents/data/fileReference_schema.json
Vendoring is the right half of that choice, not MATLAB_ONLY. A
fileReference is what makeSourceFile now writes to describe the .gef a
pyramid was built from, so it lands in any session ingested from
NDI-matlab. Python opening that session needs the definition and schema
to validate the document -- without them a session that MATLAB wrote
would fail to read for want of a file that was never copied across, which
is exactly the class of break this check exists to catch before it
reaches a user.
Byte-identical to the MATLAB copies, so the divergence check stays quiet.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
stage() printed a label and then nothing moved until the step finished.
On a directory-backed session that is fine: the reads are milliseconds.
On a cloud-backed one the cell table, the contour file and the gene list
are each a whole download before the window can appear, and a static
label for a minute is indistinguishable from a hang -- which is what a
first open of a cloud dataset actually looks like.
getFile takes an optional progress callback, called per chunk with the
running total and the Content-Length. That is the one place every
on-demand fetch passes through -- cell table, contours, gene list and
every tile all arrive by fetch_cloud_file -> getFile -- so reporting
anywhere else would cover part of the wait and not the rest.
It reaches the caller through a module-level observer rather than a
parameter, because there is nowhere to thread one through: DID fixes the
file handler's signature, counting positional parameters to choose
between the two- and three-argument forms, so a fourth would be read as
the context dict.
stage() installs the renderer itself, so every step gets the bar with no
change at the call sites and the label reads as one line:
[genepyramid] reading the cell table: 12.3 MB of 26.0 MB (47%)
Three things it deliberately does not do. It does not draw when stderr is
not a terminal, because a \r-redrawn line becomes thousands of lines in a
log or a CI capture; the stage timings still say what happened. It does
not report a total the server did not send, rather than inventing a
denominator for a chunked response. And a callback that throws is dropped
and the download continues -- reporting is never worth the bytes.
The renderer locks its own state: tiles are fetched from _TileFetcher
worker threads, so the observer is called from more than one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
The stderr bar was the wrong medium. The viewer is normally started from NDI-matlab's GEF Manager View button, which runs it as a child process with no terminal attached: the bar rendered to a stderr nobody was looking at, and reported nothing while the Dock icon bounced. So the indicator is now a small Qt window -- a stage label and a bar -- built BEFORE napari is imported, because the wait starts before there is a viewer to host it. Opening the session, the cell table and the contour file are all spent first, and on a cold cache each is a whole download. napari reuses the QApplication this creates; the high-DPI attributes are set here for that reason, since napari sets them before creating its own and an app made without them renders at the wrong scale on a Retina display. NDI_GENEPYRAMID_PROGRESS forces gui, text or off. Unset, it takes the window when one can be built, the stderr bar when stderr is a terminal, and neither otherwise. The stderr narration itself stays on in every mode: it is the record of where the time went. THE DISPLAY GUARD IS THE PART THAT MATTERS. Qt does not raise when it cannot reach a display -- it prints a line and calls abort() -- so a try/except around QApplication() catches nothing, and the first run of the test suite took pytest down with it rather than falling back. The guard checks the platform and the environment before touching Qt at all: macOS and Windows have a window server, X11 and Wayland say so in the environment, and an explicit QT_QPA_PLATFORM counts (which is how this is tested, offscreen). A headless launch now gets text or nothing, and never a dead process. Byte updates from a _TileFetcher worker are dropped rather than drawn: touching a widget off the GUI thread is a crash, not a wrong number. The window closes when the viewer is built and about to be handed over -- an indicator that outlives what it reports on makes a finished launch look stuck, and it would sit over the picture it exists to reach. 25 tests, all headless. The display guard is covered directly, because the failure it prevents cannot be caught after the fact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
Follows NDI-matlab 38aa4e8fb. A grid is a fixed FRACTION of the extent,
so the same 9x9 that gives a mouse section 20 MB tiles gave a real ferret
hemisphere 257 MB ones -- five times the budget the grid was sized for.
grid now defaults to None and is chosen from where the records actually
are, against tileBudgetBytes (50 MB by default).
The budget is on the BIGGEST tile at the finest level, not the average
one: tissue does not fill its bounding box, and on that ferret the
densest tile ran 2.5x the median, so budgeting the median misses by that
factor. gridRange (3..64) bounds the search, its upper bound a file COUNT
limit rather than a geometric one -- grid**2 files per level, all of them
written, stored and uploaded, which is a cost the density estimate cannot
see and which is therefore reported through progressFcn.
The two languages reached this from different sides. Python already
sorted on a tile-major key; MATLAB scanned every record against every
grid cell, which a data-sized grid makes untenable, and has now adopted
the Python key. Both now go further in the same direction:
- the tile, the tile-local pixel and the gene are DECODED back out of
the key rather than carried alongside it through the sort. Two
full-length arrays go through the sort where six used to.
- a level is built one band of tile rows at a time. A tile spans th
level rows and the full width of its row, so records in different
tile rows never share a tile and can be sorted and written apart. The
working set becomes the section divided by the grid -- and because
the grid is now chosen from the data, a section big enough to need
the chunking is the one that gets a fine grid to chunk with.
Neither changes a byte of what is written: a tile is a contiguous
rectangle, so global row-major order inside it IS local row-major order.
The shared conformance tile still holds both languages to it.
progressFcn(fraction, text) is new here and reports the grid it chose,
then every level and every band within a level. The pyramid is the slow
half of an ingest -- 951 s of a 971 s ferret build -- and reported only
at its ends, so minutes of it were indistinguishable from a hang.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
…cs-hdf5-format-9jn3l9
I had picked 2048 assuming the blur cost about 50 ms. Measured, with
scipy's gaussian_filter across the useful sigma range:
1024^2 38 ms (sigma 2px) 126 ms (sigma 20px)
2048^2 177 ms 442 ms
4096^2 757 ms 1815 ms
The blur is redrawn on every change of the width knob, so this sits on
the interactive path, not the loading one -- 2048 would have made the
slider lag by a third to half a second per tick. 1024 is the largest that
still redraws inside a drag.
Binning is not what costs and never was: 400,000 centroids histogram in
13 ms, once, and the result is kept. The cap is about the blur.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
…cs-hdf5-format-9jn3l9
The standing raster covers the whole section, so it goes blocky once you
zoom past it: 1024 across a 59,000-unit ferret is ~58 base pixels a
raster pixel. This re-bins just what is on screen, putting the same
number of raster pixels over a much smaller region -- full resolution for
that zoom.
ON DEMAND RATHER THAN ON CAMERA MOVE, deliberately. The blur is 38-126 ms
and camera events fire continuously through a pan, so following the
camera would stutter through every drag unless it were debounced, and a
debounce is a guess about how long someone has stopped moving. A button
makes the cost explicit and puts it where the user already knows they
have finished moving. If it turns out to be tiresome, a debounced
auto-refresh is a small addition on top of this.
Two things it has to get right, both tested with no display:
- THE RECTANGLE. napari's camera.zoom is canvas PIXELS PER WORLD UNIT,
so the visible extent is the canvas size over the zoom, centred on
camera.center. Qt reports (width, height) while napari axes are
(row, col), so width spans the columns and height the rows -- the one
thing here that is easy to get backwards and invisible whenever the
window happens to be square.
- THE FALLBACK. The canvas size is reached through private attributes
that have moved between napari versions, so it is tried rather than
assumed. A viewer that will not say leaves the blur over the whole
section and says so, because a blur over the WRONG rectangle would be
worse than a coarse one over the right rectangle.
A re-blur changes the raster's scale and translate as well as its data,
so all three are set now; data alone would leave the image stretched over
the rectangle it used to cover.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
TWO GREYS, one per language. napari defaults to its dark theme and the docks inherit it, so the viewer came out grey against every other NDI applet -- white bodies under a navy header. viewer.theme is set to "light", and each control panel gets a stylesheet built from ndi.gui.cloud_colors: the same triplets ndi.gui.cloudColors holds on the MATLAB side, so the two stay one look rather than two approximations of one. The stylesheet is applied in addAllPanels' _build rather than in each builder, so a panel added later cannot forget and come out grey among white ones. On the MATLAB side the figure and its layouts were already on c.offWhite -- but the CONTROLS were not. A uitable, a uitextarea, a uilistbox and a uieditfield each default to MATLAB's own grey chrome, and the table is the largest surface in the window, so the app read as grey with a tinted border. A paper() helper beside accent() puts them on the palette, and Browse... picks up the accent the other buttons already had. Both are defensive in the same way and for the same reason: a control that will not take a colour, or a Qt build that refuses a stylesheet, costs the colour and not the window. Losing a panel over a shade would be a poor trade. FOLLOW THE VIEW is the other half of Re-blur here, and it is DEBOUNCED rather than live. Camera events fire continuously through a pan and each redraw is 38-126 ms, so tracking them directly would stutter the whole drag. A single-shot timer restarts on every camera event and fires once movement has stopped -- the same moment a hand would have reached for the button. Off by default, and disabled outright with a reason in its tooltip if this napari will not report camera movement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
The bridge completeness gate caught my own drift. NDI-matlab 92ddd79c1 put the GEF Manager's controls on the cloud palette, and this entry still recorded 38aa4e8fb, so the check reported the file as moved -- which it was. Nothing to port. That MATLAB commit is styling only, and the Python side made the same correction independently and in its own idiom: the napari viewer sets the light theme and styles its dock panels from ndi.gui.cloud_colors, which holds the identical triplets. These are two different windows that agreed to look alike, not one window and its port. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
napariViewGEF pointed at a downloaded cloud dataset reported "no spatialGeneExpressionPyramid" for a directory that plainly had one -- the same path opened in MATLAB as an ndi.dataset.dir listed the pyramid and its levels. THE TWO LOOK IDENTICAL ON DISK. A dataset keeps its database at <path>/.ndi exactly as a session does, so ndi_session_dir opens a downloaded dataset without complaint and then finds almost nothing in it: a dataset's documents may live in LINKED SESSIONS, and only ndi_dataset.database_search follows those links -- "searches the session's database directly, then also searches linked sessions", as it says. The session reader looks in the dataset's own database and stops. Nothing errors. The path is right, the data is there, and the reader is one level too shallow. The dataset reading is tried first and kept when it finds pyramids. A plain session opened as a dataset has no linked sessions and answers the same as before, but that is not relied on: if the dataset reading raises, or finds nothing where the session reading finds something, the session reading wins. Neither working is the only error, and it names both. _asDataset and _asSession are named functions rather than inline imports so the tests can stand in for them -- what _open_session does AROUND the two readings is not about how either is constructed. For the session one there is a second reason: ndi.session.dir resolves to the CLASS, not the module, because the package attribute shadows the submodule, so patching it where it is imported does not work. That cost a test run to find and is worth a line rather than rediscovering. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
Opening a downloaded dataset in the viewer got past "no
spatialGeneExpressionPyramid" and then failed on the first tile:
RuntimeError: Cannot fetch 'tile.bin_262' from thread
ThreadPoolExecutor-1_0: this session cannot be reopened for another
thread ...
That is the tile fetcher's own guard, and it was right to fire. NDI's
database belongs to the thread that opened it and dask runs blocks on a
pool, so the fetcher builds a per-thread handle by reopening whatever it
was given AT ITS OWN PATH. ndi_session_dir exposes `path` and could be
reopened. ndi_dataset_dir kept the same value privately as `_path`, so
_reopener found nothing, _ensurePool returned None, and the fetcher
could only explain itself and stop.
MATLAB's ndi.dataset.dir has declared a public `path` property since it
was written ("the file path of the session"), so this was a parity gap
rather than a design choice: this side simply never exposed the _path it
was already keeping. A directory-backed dataset that cannot say which
directory is a hole regardless of who asks.
The reopened handle resolves tiles correctly because
ndi_dataset.database_openbinarydoc already dispatches to the doc's owning
session, member sessions included -- the case its own docstring says
would otherwise raise FileNotFoundError "even when the file is really
there".
Found by fixing the layer above it. Opening a downloaded dataset AS a
dataset (71df5c9) was necessary and not sufficient: it moved the failure
from "no pyramid" to the first tile, because it changed the type the
fetcher was handed to one that could not be reopened.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
…cs-hdf5-format-9jn3l9
NDI-matlab cf0983b5c made the View dialog non-modal and gave the app an onClose. Nothing to port -- the Python side draws no dialog at all; View builds a command line and runs napariViewGEF, which is already its own process and has never blocked anything -- but the hash has to follow the file or the completeness gate reports drift, as it correctly did the last time I moved this file without bumping it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
…lazy
TWO SURFACES, and I had confused them. The ask was for the control
PANELS to be on the NDI Cloud palette like every other applet; I set
napari's whole theme to "light", which did that and took the CANVAS with
it. Imaging data is read against a dark ground, and the canvas is not a
panel. The theme is left alone now; the panels are still styled
individually by controls.applyCloudStyle, which does not touch the
canvas.
And the launch reporting was lying about where the time goes. One stage
wrapped both layerSpec and add_image under the label "building the
pyramid ladder (lazy)". Only the first of those is lazy:
add_image -> Image._post_init -> refresh -> set_view_slice
-> _level_materializer(thumbnail_level) -> np.asarray(data[level])
napari builds its thumbnail by materialising the WHOLE COARSEST LEVEL. On
a real section that is gigabytes of tiles, and on a downloaded dataset it
is gigabytes of FETCHING -- minutes, behind a label claiming nothing was
being read. A ferret hemisphere sat there long enough to be reported as a
hang, and it was not hung; it was working, silently, on the step the
label said was free.
Split in two, and the second one named for what it does. That does not
make it fast -- the coarsest level really is that big, because binning
merges pixels while the genes in them stay distinct -- but it makes the
wait legible instead of looking like a stall. Making it actually fast
needs a small gene-summed overview raster stored at build time, which is
a rebuild and a separate decision.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
"The GEF manager window and the viewer window are mostly gray" meant the GEF Manager and the VIEW-IN-NAPARI LAUNCHER DIALOG -- both MATLAB uifigures, and both already on the cloud palette since NDI-matlab 92ddd79c1 put their uitable, uitextarea, uilistbox and uieditfield on c.white beside the layouts that were already c.offWhite. I had read "the viewer window" as napari and changed two things there that nobody asked for. The theme went back in 8f73cf0, which returned the canvas to dark. This takes the other one: the dock panels are no longer restyled. Removed rather than left dormant, because on reflection it was also the wrong look. napari's own layer list and layer controls stay in napari's theme, so white panels beside them would be two palettes in one window -- a worse result than either alone, and one that would read as a bug rather than a decision. The MATLAB side is unaffected: those windows are ours, sit beside the other NDI applets, and should match them. The napari viewer is a third-party window we borrow, and it should look like itself. One line to put back if it turns out to be wanted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
The launch window sat on "importing napari" for minutes while the label underneath had moved on several steps -- so a launch that was working read as one that was stuck, on the wrong step. The label was never wrong. setText only QUEUES a repaint, and processEvents services what is queued at that instant. The step that follows then blocks the thread with no event loop running, so any paint the window needs AFTER that point -- it gets exposed, the compositor asks again, the OS marks the app unresponsive -- is never serviced. The window freezes on whichever label was current when the first slow step began, and on a cold launch that is the napari import. repaint() paints synchronously, before returning, so each label is on the glass before its step blocks. The byte counter gets the same treatment, for the same reason: a byte count that only lands once the transfer has finished has reported nothing. This does not animate, and cannot: without an event loop nothing moves between steps. What it buys is that each step is legible WHILE it runs, rather than one step being legible for all of them. Making the window live through a blocking step would mean running the work off the main thread, which napari's viewer cannot be. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
"lazy" is a programmer's word for it, and it was not the only one. The
labels named our implementation and napari's API rather than what the
user is waiting for:
importing napari -> starting napari
building the pyramid ladder (lazy) -> preparing the image levels
(read on demand)
drawing the overview (napari reads -> drawing the overview (reads
the coarsest level) the whole lowest-resolution
level)
napari add_shapes -> drawing the cell outlines
napari add_points -> drawing the cell centroids
"add_shapes" and "add_points" were method names showing through. "ladder"
and "coarsest level" are how the pyramid is built, not what is happening
to the picture.
The parentheticals stay, because they are the two places where knowing
WHY the wait is what it is changes what the user does: "read on demand"
says the tiles have not been fetched yet, and "reads the whole
lowest-resolution level" says why the step after it is the long one. The
outline note loses "triangulates" for "cut into triangles", which is the
same fact without the word.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
A gene's counts are a handful of pixels scattered over a section. Drawn honestly they are dust, and a reader looking for where a gene is expressed cannot see a pattern in it. blurLevels spreads each pixel over its neighbourhood so the pattern reads. Two things make it more than a call to gaussian_filter: THE WIDTH IS A DISTANCE, NOT A PIXEL COUNT. Each level of the ladder is binned differently, so a fixed number of level-pixels would be a different physical distance on each -- the blur would visibly change width as the viewer switched levels while zooming, which is exactly what a multiscale ladder exists to avoid. Dividing by each level's bin size makes one setting mean one distance everywhere. The knob itself is in micrometres, since that is the unit a reader thinks in; basePixelUm reads base_pixel_size_x to convert, falling back to the Stereo-seq 0.5 um pitch rather than dividing by zero on a document that does not say. BLURRED PER TILE, WITH A HALO. Levels are dask arrays whose blocks are tiles, and a plain per-block filter sees each tile's edge as the end of the world -- leaving a seam at every boundary, in a grid, which reads as an artefact of the data rather than of the drawing. map_overlap gives each block truncate*sigma of its neighbours, which is the kernel's reach, so the result matches blurring the level whole: max error 0.0 against scipy over sixteen blocks. The kernel is unity, so widening spreads counts rather than adding to them and width stays independent of brightness. The peak does fall a long way, so each re-blur resets contrast limits; without that a widened layer goes black. The unblurred ladder is kept per layer and every change re-blurs from it, so the knob is absolute rather than compounding. This replaces the centroid-density raster, which blurred the wrong thing: the ask was the gene data, not the cell dots or the contours. Its 360 lines and its tests go. Rotation panel: the slider shared a row with the angle box and the Reset button, which left it a stub too short to aim with -- a dock panel is narrow and those two take a fixed width out of it whatever is left over. It now gets the full width of a row beneath them. The box's own - and + buttons step five degrees, a turn you can see, where half a degree took ten clicks to show anything; finer angles are still typed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
The band slider shared a row with its label, and it carries its own value labels at each end of the track -- so in a dock this narrow what was left after "Abundance band" was the two numbers and no track between them. Unreachable rather than merely cramped. The label is now the line above and the slider has the row to itself, which is the same lesson the rotation slider taught yesterday. The gene list came up four rows tall, wedged between the band above and the blur below. Finding a gene among 26,444 through a porthole is not a control. A dock panel is only as tall as it is asked to be, so it is now asked: a 260px minimum on the list, which the panel's stretch already hands any further space to. Both are checked on real widgets built offscreen, since an arrangement is a property of a layout and stubbing the widgets would only test the stubs. They fail on the old layout and skip where there is no Qt binding -- which is CI today. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
"Binary file 'tile.bin_280' not found for document ..." is what an authentication failure looks like from the viewer: the tile is in the document's file list and has an ndic:// location, so it exists -- NDI simply could not authenticate to fetch it. The message underneath said "No valid token and no credentials available", which is equally true of never having logged in and of a token that lapsed while the window was open, and those want opposite responses. authenticate() now names which of its three steps came up empty. The expired case in particular says so, says when, and says the thing a re-login does not reveal: the token lives in os.environ, so it belongs to ONE PROCESS. Logging in from a shell -- or from MATLAB, whose vault Python cannot read -- leaves a separately launched viewer exactly as unauthenticated as it was. The remedy named in every case is credentials authenticate can reach, because with those it renews the token itself and the first step stops mattering. credentialReport() walks the same three steps and reports each, to be run IN THE ENVIRONMENT THAT FAILED. It reports the token by presence and expiry, never by value, and prints no password, so it can be pasted into an issue. Why this surfaces as scattered tiles rather than a blank picture: tiles already in DID's local cache need no credentials, so a mostly-downloaded section draws and the gaps are wherever the cache is cold. Blur widens that -- map_overlap gives each visible tile its eight neighbours -- so turning it up reaches tiles a still view never asked for. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
The rotation panel was built with the layers that existed when the viewer opened and never asked again. Tick a gene while the section is turned and its layer arrives SQUARE: the gene then sits beside the anatomy rather than on it, which is a wrong picture rather than an unrotated one, and nothing on screen says which of the two layers is the one to believe. Two halves, and both were missing. The panel now takes the viewer's own layer list at the moment it applies -- "what is on screen" is a question with a current answer, not one settled at launch -- so a later rotation reaches everything added since. And it connects to the layer list's inserted event, so a layer added while a rotation is in force is turned as it arrives. THE MATRIX IS HANDED OVER, NOT RECOMPUTED. The pivot is re-read between interactions, so rebuilding the affine from the angle would pivot the new layer about wherever the view centre had got to; two layers rotated by equal angles about different pivots do not line up. The one in force is kept and applied verbatim, which is also why a square picture leaves an arriving layer alone rather than stamping an identity over an affine some caller set for its own reasons. Six tests, four of which fail on the old fixed-list behaviour. The insert hook is wrapped because a headless caller composing layers by hand has no emitter, and a control that raises on construction costs the whole picture. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
A VIEWER OUTLIVES ITS TOKEN. Someone reading a section works for hours; the token is good for rather less. What running out looks like from here is not a login prompt but MISSING DATA -- tiles already cached keep drawing, and the ones that were never fetched fail one at a time as the reader pans, which reads as a broken pyramid. The token lives in this process's environment, so nothing outside it can reach in: the remedy was to quit, log in elsewhere, and open the section again. An NDI Cloud dock now shows the time left and keeps it current, warns at fifteen minutes rather than at zero, and takes an email and password to renew in place. Signing in writes the token where every later fetch reads it -- including the ones on the tile threads, since a fetch builds its client per call rather than holding one, so a renewal takes effect on the next tile with nothing to restart. THE PASSWORD IS NOT KEPT. It goes to the login call and nothing else: not to a field that stays filled in a dock someone is screen-sharing (it is cleared whether the attempt succeeded or failed -- the failed one is the likeliest typo), not to a file, and to the profile only when the checkbox asks. That checkbox exists because credentials saved once and wrong are exactly how someone ends up here, so it CORRECTS the entry that already carries the email rather than adding a second one beside the wrong one. The panel is offered only when a token is in the environment, which is the same signal the MATLAB launcher uses and means the parent process had been talking to the cloud. A local pyramid gets nothing: a login control on a window that needs no login implies the picture might be waiting on something. Also in auth: tokenSecondsRemaining and tokenStatusLine, which say "signed in, 30 min left" or "EXPIRED 20 min ago". An opaque token reads as unknown rather than expired -- calling it expired would sign out someone who is signed in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
#297 added ndi_dataset_dir.path independently, as the fix for #295 -- the issue this branch's copy was filed from. Same property, same reason; the only conflict was which docstring says it. Main's is kept: it cites the issue, names the failure it produces three layers away ("Cannot fetch 'tile.bin_262' from thread ThreadPoolExecutor-1_0"), and records that the setter is protected on the MATLAB side, which this one matches by being read-only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
…them Two panels, both acting on the "gene: XXX" layers and NOTHING ELSE. The base section and the cell overlays are deliberately excluded: the anatomy underneath is what the genes are read against, and dimming it to make a gene stand out takes away the frame that makes standing out mean anything. GENE APPEARANCE. Gene layers arrive with limits set from their own data, which is right for reading one gene and wrong for comparing several -- two genes an order of magnitude apart get drawn as though they were the same brightness, and the picture says less than the data does. One contrast and one gamma now move them together. Contrast narrows each layer's window about its FLOOR rather than its centre, because these layers start at zero and blend additively; raising the floor would punch holes in the background where a gene is merely absent. Above 1 brightens and separates a sparse gene, since the dim end of the range is where all its counts are. Each layer keeps its OWN reference window -- the one factor applied to windows that differ by an order of magnitude is the point, where a single shared window would black out the quiet gene -- and every change recomputes from that reference rather than from the current limits. So three nudges up and one back down returns you to where you started, and the number on the control is the whole state. A gene ticked later picks up the settings in force, or the comparison the panel exists for would be broken by the act of adding to it. GENE TOUR. One button walks the gene layers, pulsing each for five seconds -- gamma swept down and back once a second, twenty frames a second -- with the gene's name in 24pt at the bottom right of the canvas. A legend of six colours against a section this dense is a legend nobody can read, and a demo is usually watched on somebody else's screen, so the picture says which is which instead. Nothing blocks: the sweep is a timer stepping a small state machine, so the viewer stays live and a tour can be stopped mid-gene. What it touches it puts back -- each gamma is recorded when the tour arrives and restored when it leaves, so stopping halfway cannot strand a gene at a value nobody chose. It pulses around the gamma the appearance panel set rather than around 1, so the two controls agree. The tour clock is a STEP COUNT, not an accumulated float. Adding 0.05 a hundred times lands near five seconds but not on it, and the gene boundary is a comparison against exactly that -- so it would fall a frame early or late depending on the interval, which is the kind of bug that only appears on somebody else's machine. Gamma bottoms out at 0.01 rather than 0: napari uses it as a divisor and rejects zero, and the difference is not visible. 43 tests on the two panels plus 3 on the wiring, all offline -- the tour's own step function is driven by hand rather than waiting five seconds a gene. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
The pulse tour leaves every gene on screen and sweeps one gene's gamma, so it answers "where is THIS one, among the others" -- the comparison is the point and the sweep is the pointer. That is not the only question worth asking. Six layers blending additively are a colour nobody can decompose by eye, and "what does this one look like by itself" needs the others out of the way. So: a second button, five seconds a gene, hiding the rest. Same clock, same canvas name overlay, same state machine -- the modes differ only in what they touch between gene boundaries, and solo touches nothing there at all: the switch happens once, when the tour arrives. THE BASE SECTION IS NOT HIDDEN. Only the gene layers are, which is the whole difference between showing one gene and showing one layer -- a lone gene on black has nothing to be read against, and the anatomy underneath is why the picture means anything. IT RESTORES WHAT WAS, not all-on. A reader who had already turned a gene off did that on purpose, and a tour is not a request to undo it. Stopping halfway restores too: five of six genes left hidden is a worse state than the tour started in. ONE AT A TIME. Starting either mode stops the other first, so the two can never both be moving the same layers -- and stopping first is what puts back the gamma the pulse was mid-sweep on, which would otherwise stay stranded for the whole of the next tour. 17 more tests. Two deliberate regressions confirm they bite: restoring all-on instead of what was, and switching modes without stopping. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
Was magenta, green, cyan, yellow, red, blue -- an order with nothing behind it. Genes are coloured in the order they are ticked, so this tuple's order is the order a reader watches them arrive, and a spectral run is one they can hold in their head and read back off the picture where an arbitrary one is six colours to memorise. Now red, yellow, green, cyan, blue, magenta: long wavelength to short. Magenta is not spectral at all -- it closes the circle, which is where the eye expects it after blue. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
…rter stack FOUR THINGS, all in the viewer. --gene-layers. A new option that opens named genes as their OWN additive layers on top of the base image, each in a named colormap: "HPCAL1:green,RORB:blue,FEZF2:bop orange". A DIFFERENT REQUEST from --genes, which filters the base image down to a subset and leaves one picture -- this keeps "All genes" and adds a layer beside it, which is what makes several genes comparable. Both can be given together. A colour is optional and means "take the next one from the cycle", so a caller can name the ones that matter. A named colour does NOT advance that cycle: otherwise the first hand-ticked gene's colour would depend on how many were preset, a surprise with no upside. Colormap names may contain spaces -- napari ships "bop orange" -- so only the ends of each field are trimmed. Genes are ticked through the panel's own checkbox rather than added behind its back, or unticking a preset gene would do nothing at all; one not in this pyramid's list is named on the panel, because a demo that quietly opens four of six layers looks like the viewer working. NDI-matlab's Ferret demo preset builds this string. THE CLOUD SIGN-IN IS BEHIND A BUTTON. One row now: the button and the time left. Three permanently docked fields said a reader signs in often; they do it once a day at most. The dialog offers the SAVED PROFILES first -- someone whose profile works should not retype a password to get past an expired token, and someone whose profile does not work can see which one is being used, which is half the diagnosis. Picking one fills the email; picking one with a stored password signs in with it; a typed password still wins, which is the point of arriving there. Built and shown are split, so a test can drive the fields and the Sign in button with no event loop to escape from. That split is not theoretical: the old docked-fields tests clicked a button that now opens a MODAL dialog, and the suite hung until it was killed. CONTRAST AND GAMMA ARE SLIDERS ONLY. They are dragged, not typed, and two spin boxes were most of two panels' height. The value is still shown beside each slider, so a setting worth writing down can be read off. THE GENE LIST IS LAST, so it sits at the bottom. It is the tallest panel by a long way and napari gives leftover height to whatever is furthest down, so anything below it gets squeezed while the list takes the room. The appearance and tour panels read above it, which is the order they are used in; the cloud row moves up with the other whole-session controls, since it is about the connection rather than the picture. THE BLUR BOX WAITS 1.5 SECONDS. A spin box emits on every keystroke, so typing "50" asked for 5 um first -- and re-blurring is a whole-ladder map_overlap that takes long enough to see, so the picture visibly redrew at the wrong width before the second digit landed. Enter or clicking away fires at once rather than sitting out the wait. The band slider keeps its 180ms: dragging a slider wants to be followed, while a number being typed is not a request until it is finished. Also fixed while testing: the "not in this pyramid" warning was set before _report(), which rewrites the same label -- so the warning was one nobody could see. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
A profile whose secret lives in a store this process is not reading looks exactly like a profile with no secret at all, and the sign-in dialog reported it as the latter. That is not a rare corner: profile._detect_backend picks keyring whenever keyring merely imports, while MATLAB writes its secret to the AES file -- so a machine with both looks in the Keychain, finds nothing, and tells the reader their password is missing. They then re-save a password that was already saved, into a store nothing is reading. cloudProfileChoices now returns a note alongside each row and the chooser shows it: "password saved", "no password in the keyring store", or "password unreadable from the aes store (<reason>)". Picking a profile whose password is not usable puts that reason on screen rather than only moving the focus, because "not readable from here" and "you never saved one" are different problems and only one of them is solved by typing it again. Diagnosis only. Whether the backend preference itself is wrong is a separate question, waiting on what the store actually reports on the machine where this was seen. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
The cloud row is the only panel about whether the picture can still be READ rather than about how it is drawn, and its clock is the thing worth catching before it bites: a reader who sees "12 min left" at the top of the stack can sign in at a convenient moment, where the same line at the bottom is found afterwards, by way of a tile that will not load. So it goes first. The gene list stays last, where the leftover height belongs. The wordmark is copied from NDI-matlab (+ndi/+cloud/+ui/resources/images/ndi_logo.png) rather than referenced, because the viewer runs from a Python install that need not have NDI-matlab on disk at all -- a downloaded dataset opened by someone who has never run MATLAB is the normal case. A test pins the md5, so the two repos cannot drift into showing different marks for the same product. ON A WHITE CARD, because it is dark navy on transparency and napari's docks follow the viewer's theme: on the dark one the lettering is simply not there. That is how the brand is used elsewhere anyway. Scaled to twice its display width at a device pixel ratio of 2, so it stays crisp on the retina screen a demo is usually given from. A missing asset costs nothing: cloudLogoPath returns None, the label is skipped, and the panel keeps its dock name in text. Nothing here changes the secrets backend. The profile that would not load turned out to be a machine that had never had the password set -- the AES file was simply absent, which rules out the keyring/aes mismatch I had guessed at rather than confirming it. Setting it in MATLAB's profile editor made both sides read it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
ONE EMAIL CAN OWN SEVERAL ACCOUNTS. A dev profile and a prod profile under the same address is the normal arrangement, not a mistake to be tidied -- and saveCloudProfile took the FIRST entry matching the email, wrote the password into it, and then called set_default on it. Both halves are wrong. Writing a prod password into the dev entry is silent, and re-pointing the default while doing it is worse: everything afterwards authenticates as the wrong account, against the wrong server, because CloudConfig.from_env resolves the API URL from the profile's Stage. That failure surfaces as a 404 on the dataset, which reads as a permissions problem rather than as a wrong-server one -- the same trap this branch already documents for the dev/prod token short-circuit. So it never guesses. In order: the uid the caller names -- and the dialog now passes the profile the reader actually picked, which settles it outright -- then the CURRENT profile, then the DEFAULT, when either matches the email, then a match that is the only one for that email. Several matches and no steer raises, naming the count and what to do about it. An ambiguous save is the one case where doing nothing and saying so beats acting. AND THE DEFAULT IS NOT MOVED for an existing profile. Saving a password is not a request to switch accounts. A profile this creates is still made the default, because a store that had nothing for that email has no other candidate to displace. Ten tests, including the one this came from: two entries sharing an email with the dev one listed first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9
stevevanhooser
deleted the
claude/spatial-genomics-hdf5-format-9jn3l9
branch
September 9, 2026 20:28
stevevanhooser
pushed a commit
that referenced
this pull request
Sep 9, 2026
Main #291 (spatial-genomics-hdf5-format merge) also added an entry for +ndi/+util/hdf5Contents.m under the porting_deferred: heading, with a more accurate decision_log about why it is not ported yet (MATLAB-console convenience for interactive ingest). The merge in 5139ded pulled that in, and my earlier matlab_only entry from c248b09 is now redundant. Keep main's; drop mine. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N67xH9BejMzZp4GNAq8m7w
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replaces #290, which was opened on a branch named
…-9jn3l9-viewer. The bridge CI job checks out NDI-matlab atgithub.head_refand falls back tomainwhen that branch does not exist, so the-viewersuffix would have paired this PR against NDI-matlabmain— whereGEFManager.m,hdf5Contents.mand the tile helpers do not exist yet, failing every path and sync-hash check. Same commits, branch renamed to match VH-Lab/NDI-matlab#977, now merged.The Python half of the spatial-genomics viewer work. Five groups: the napari viewer itself, the controls that make it usable during a talk, its launch cost, the cloud handoff from NDI-matlab, and a file-naming convention we were on the wrong side of.
The viewer
napariViewGEFopens a pyramid as a lazy multiscale layer with cell centroids and outlines. This PR adds the panels that make it usable rather than just openable:leidenandsubclass_nn_columnagree on 99%+ of cells in the opossum data; rather than picking one, both are shown and the agreeing pair is named.A panel that throws no longer takes the window with it — each is built inside a guard that prints the reason and leaves the image alone, because a broken control panel is not a reason to lose the picture you waited for.
Controls for reading a section, and for showing one
Rotation. Sections are not mounted square. An arbitrary angle about the centre of the current view, turning the image, the centroids and the outlines together — one shared affine, since a control that could slide them apart is a control that can produce a wrong picture. The pivot is captured when the interaction starts, so panning mid-drag cannot make the picture jump. It governs layers added later too: the panel reads the viewer's own layer list at the moment it applies, and hooks the
insertedevent so a gene ticked while the section is turned arrives turned. The matrix is handed over rather than recomputed — equal angles about different pivots do not line up.Blur. Gene expression is a handful of counts in scattered pixels; drawn honestly it is dust. A µm-denominated Gaussian spreads it into a pattern. The width is a distance, divided by each level's bin size, so it does not change as you zoom through the ladder. Blurred per tile with a
truncate*sigmahalo viamap_overlap, so there is no seam grid — measured max error 0.0 against blurring whole. Unity kernel, so width and brightness stay separate; the peak falls a long way, so contrast limits are reset each time. The unblurred ladder is kept per layer, so the knob is absolute rather than compounding. The box waits 1.5s after you stop typing: it emits on every keystroke, so "50" used to render at 5 on the way.Gene appearance. One contrast and one gamma across every
gene:layer at once — gene layers arrive with limits set from their own data, which is right for reading one gene and wrong for comparing several. Contrast narrows each layer's window about its floor, because these layers start at zero and blend additively and raising the floor would punch holes where a gene is merely absent. Each layer keeps its own reference window, which is what lets one knob serve genes an order of magnitude apart.Gene tour, two modes. Pulse leaves everything on screen and sweeps one gene's gamma down and back, once a second for five seconds — "where is THIS one, among the others". Solo hides the other genes and shows one at a time — "what does this one look like by itself", which additive blending otherwise makes hard. Both name the gene in 24pt at the bottom right of the canvas, because a legend of six colours against a section this dense is a legend nobody can read. Neither blocks: one timer stepping a state machine, so the viewer stays live and a tour can be stopped mid-gene. What they touch they put back — and solo restores what was, not all-on, since a gene the reader had already hidden was hidden on purpose.
The base "All genes" layer and the cell overlays are excluded from all of these. The section underneath is the anatomy the genes are read against; dimming it to make a gene stand out would take away the frame that makes standing out mean anything.
--gene-layers, and the Ferret demoA new option opening named genes as their own additive layers on top of the base image, each in a named colormap:
A different request from
--genes, which filters the base image down to a subset and leaves one picture: this keeps "All genes" and adds a layer beside it, which is what makes several genes comparable. Both can be given together. A colour is optional and means "take the next from the cycle"; a named colour does not advance that cycle, or the first hand-ticked gene's colour would depend on how many were preset. Colormap names may contain spaces — napari shipsbop orange— so only the ends of each field are trimmed and the whole value is shell-quoted as one argument.Genes are ticked through the panel's own checkbox rather than added behind its back, or unticking a preset gene would do nothing at all. One not in this pyramid's list is named on the panel: a demo that quietly opens four of six layers looks like the viewer working.
The Ferret demo preset that builds this string lives in NDI-matlab's
GEFManager, not here — the viewer takes genes and colours as ordinary options, so a demo is a preset command rather than a mode the viewer has to know about.Launch cost
~27s → ~5s before the window appears, measured at opossum scale (80,369 cells, 843,189 contour vertices). Four things, none of them clever:
list(zip(row, col))per polygonnp.stackper cellAll replaced with whole-array operations; the contour file is now read into one
(total, 2)array that isnp.splitinto views. Stage timing goes to stderr so the next regression is visible rather than felt (NDI_GENEPYRAMID_QUIET=1silences it).A launch window replaces the stderr-only progress, since stderr is not where anyone was watching. Each label is painted with
repaint()before its step blocks —setTextonly queues a repaint, and a blocking step with no event loop leaves the window frozen on the last label, which reads as a hang. The labels say what the user is waiting for rather than what the code is doing: "starting napari", "preparing the image levels (read on demand)", "drawing the overview (reads the whole lowest-resolution level)". That last one is the long step of a cold launch, and naming it does not make it fast — it makes it legible.The cloud handoff
Opening a downloaded dataset failed at the first tile fetch. Four bugs in a row, each hiding the next:
cryptographywas undeclared.ndi.cloud.profiledocuments AES-128/CBC as its default secrets backend; without that package the detector fell through to the in-memory backend, which persists nothing, and said so only through a logger warning. A default that nothing installs is not a default, so it is now a dependency rather than an extra.jsonencodewrites a 1×1 struct as a JSON object and a 1×N struct array as an array, so a file with exactly one profile — what every new user has — was iterated as a list, yielding the dict's keys. Zero profiles, no error.authenticate()never read the profile. It went token →NDI_CLOUD_USERNAME/PASSWORD→ raise, with no step in between, while MATLAB'sauthenticate.mhas hadauthenticatedWithSecretall along. Every implicit login had this hole, most visibly thendic://handler fetching pyramid tiles.matlab.lang.makeValidName, which deletes whitespace and capitalises the following letter. MATLAB wroteNDICloud41269713c8f30937_…; Python looked forNDI_Cloud_41269713c8f30937_…. Neither could read the other's secrets in either direction. Reads fall back to the old name so nothing already saved is lost.Then a fifth, found while confirming the fourth:
login()exportsNDI_CLOUD_TOKEN, so every config built afterwards short-circuits atauthenticate()'s first step and keeps the prod default URL. Adevprofile logged in against dev and issued everything else against prod, where the dataset 404s — indistinguishable from a permissions error, since a dataset that is not on this server is exactly as absent as one you cannot see. Stage now resolves inCloudConfig.from_env, matchingprofile.switchProfile'ssetenv('CLOUD_API_ENVIRONMENT', prof.Stage).A viewer outlives its token. Someone reading a section works for hours; the token is good for rather less, and it lives in
os.environ— so a login in another shell does not reach a window that is already open. What running out looks like from the viewer is not a login prompt but missing data: tiles already cached keep drawing, and the ones never fetched fail one at a time as the reader pans. So:authenticate()now names which of its three steps came up empty, and an expired token says so, says when, and says that logging in elsewhere will not reach this process.credentialReport()walks the same three steps and reports each, to be run in the environment that failed. It reports the token by presence and expiry, never by value.Opening a downloaded dataset
Two more, found on the ferret data:
session. Identical on disk, but a dataset's documents can live in linked sessions, which onlyndi_dataset.database_searchfollows — so the pyramid document was invisible and the viewer reported no pyramid at all. The CLI now tries dataset first and keeps it if any pyramid is found._TileFetcherbuilds a per-thread session handle by reopening whatever it was given at its own path, andndi_dataset_dirhad no publicpath— the property MATLAB has always had. Added here and, independently, on main as the fix for Public MATLAB properties kept private in Python: audit, and gate it like sync-hash drift #295; main's docstring is the one kept.Tile files count from one
did.document/addFileSeriestakes "ONE-BASED member numbers" and rejects anything else.makePyramidnamed its first tiletile.bin_0, putting every pyramid one step outside the convention its storage layer assumes.This is not cosmetic. NDI-matlab's uploader walks a legacy
NAME#series asNAME1, NAME2, …and treats the first absent name as the end of the series, because that gap was its only termination signal. One real pyramid reached the cloud as 24 of its 450 tile files, with the upload reporting success — documents and binaries travel separately, so the dataset downloaded and opened looking complete and failed at the first tile read instead.The tile index stays zero-based: it is a grid position, and
index_orderdefines it asrow*tile_columns + column. Only the suffix moves. The origin is recorded on the document rather than inferred — a sparse pyramid whose first tile is empty has no_0under either convention, and a wrong guess shifts every tile one cell along its row, which is an image that is quietly wrong rather than one that is missing. Pyramids already built record nothing, read as 0, and keep opening.Both directions are covered:
tileFileNamemaps index → name,tileIndexFromNamemaps back for the two readers that walk stored names and work out (row, column) from the suffix.Building the pyramid
makePyramidalso got the memory work: ~90 GB → ~13 GB on 771,217,356 records, from three changes — records kept in their native type through the sort rather than promoted to double, the tile/pixel/gene triple decoded back out of the sort key instead of carried alongside it, and levels built a band of tile-rows at a time. The tile grid is now sized from the data against a per-tile byte budget rather than pinned at 9×9. Tiles come out byte-identical across 300 randomized geometries.Merging main
Three of main's changes met this branch head-on and are worth naming, since the resolutions were judgement calls rather than mechanical:
levelArrays— main replaced this branch's eager tile-path resolution with a lazy_TileFetcher, which fixes the exact cloud case the eager version's own comment called a real limit (database_openbinarydocfetches, so resolving every path up front downloads the whole pyramid before anything is drawn). Main's version taken wholesale; only the tile-name lookup re-applied on top.profile.py— main refactored_write_secrets_fileonto a new_write_owner_onlyhelper while this branch added neighbouring functions. Pure adjacency; both kept.ndi_dataset_dir.path— added on both sides independently, as described above.Two bridge-metadata mistakes of mine that main's newer checks caught, both real:
not_applicable/not_yet_portedare retired spellings (nowmatlab_onlyandporting_deferred), and five entries recordedgit hash-objectoutput, which is a blob — the field wants a commit, orgit log <hash>..HEAD -- <path>has nothing to walk from.Tests
Roughly 300 new tests, all offline — no display, no network, no napari. Panel arrangement is checked on real widgets built offscreen and skipped where there is no Qt binding, since stubbing the widgets would only test the stubs. The ones worth naming:
scipyover a 16-block dask array: max error 0.0, which is the whole point of the halo.<16hex>_<16hex>UID, MATLAB field name — and reads the password back through the public API, so it covers the whole handoff rather than one field.abundanceBandis pinned on the case that surprised me while writing it: a loud gene's midpoint sits below a small top cut.CI is green: the full suite plus the cross-language symmetry job.
🤖 Generated with Claude Code
https://claude.ai/code/session_01YHxsZ56GZsZ3hJ98TJ81t9