Skip to content

feat: Updates for 10x Genomics Atera bundles - #426

Open
stephenwilliams22 wants to merge 9 commits into
scverse:mainfrom
10XGenomics:atera-reader
Open

stephenwilliams22 wants to merge 9 commits into
scverse:mainfrom
10XGenomics:atera-reader

Conversation

@stephenwilliams22

@stephenwilliams22 stephenwilliams22 commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Adds a native reader for 10x Genomics Atera datasets: cell-by-gene expression table, cell/nucleus segmentation labels and boundary polygons, per-transcript locations, and morphology images.

  • Table (cell_feature_matrix.zarr.zip/csc_cell_feature_matrix.zarr.zip) is read directly with anndata's own zarr IO (anndata.io.read_elem/ anndata.io.sparse_dataset), with X lazily backed by dask.
  • Cell/nucleus labels are lazily backed by dask; boundary polygons are read via shapely.from_ragged_array, with an optional tiled/pyramided partial-read path (read_cell_boundaries) for large bundles.
  • Transcripts are read lazily per spatial tile via dask-delayed.
  • Morphology images are read via tifffile's native OME-TIFF tile grid, avoiding materializing full-resolution planes.
  • An example notebook with a finalized tiny atera output bundle is provided for testing. This can be found in the atera_example.ipynb for testing.

Enables massive datasets on a basic laptop (2M cells, 11 billion transcripts). Basic analysis workflow (loading, plotting, HVG, PCA, normalization, clustering, etc) max memory usage was 7GB on a new mac in ~15min.

Adds a native reader for 10x Genomics Atera datasets: cell-by-gene
expression table, cell/nucleus segmentation labels and boundary
polygons, per-transcript locations, and morphology images.

- Table (`cell_feature_matrix.zarr.zip`/`csc_cell_feature_matrix.zarr.zip`)
  is read directly with `anndata`'s own zarr IO (`anndata.io.read_elem`/
  `anndata.io.sparse_dataset`), with `X` lazily backed by dask.
- Cell/nucleus labels are lazily backed by dask; boundary polygons are
  read via `shapely.from_ragged_array`, with an optional tiled/pyramided
  partial-read path (`read_cell_boundaries`) for large bundles.
- Transcripts are read lazily per spatial tile via dask-delayed.
- Morphology images are read via `tifffile`'s native OME-TIFF tile grid,
  avoiding materializing full-resolution planes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@review-notebook-app

Copy link
Copy Markdown

Check out this pull request on  ReviewNB

See visual diffs & provide feedback on Jupyter Notebooks.


Powered by ReviewNB

@codecov-commenter

codecov-commenter commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 23.29317% with 382 lines in your changes missing coverage. Please review.
✅ Project coverage is 60.17%. Comparing base (261578f) to head (91ed96d).

Files with missing lines Patch % Lines
src/spatialdata_io/readers/_atera_common.py 14.78% 242 Missing ⚠️
src/spatialdata_io/readers/atera.py 15.66% 140 Missing ⚠️

❗ There is a different number of reports uploaded between BASE (261578f) and HEAD (91ed96d). Click for more details.

HEAD has 5 uploads less than BASE
Flag BASE (261578f) HEAD (91ed96d)
8 3
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #426      +/-   ##
==========================================
- Coverage   65.79%   60.17%   -5.63%     
==========================================
  Files          26       28       +2     
  Lines        3263     3761     +498     
==========================================
+ Hits         2147     2263     +116     
- Misses       1116     1498     +382     
Files with missing lines Coverage Δ
src/spatialdata_io/__init__.py 100.00% <ø> (ø)
src/spatialdata_io/_constants/_constants.py 100.00% <100.00%> (ø)
src/spatialdata_io/readers/atera.py 15.66% <15.66%> (ø)
src/spatialdata_io/readers/_atera_common.py 14.78% <14.78%> (ø)
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@Zethson

Zethson commented Oct 2, 2026

Copy link
Copy Markdown
Member

@stephenwilliams22 could you please always ensure that the pre-commit checks are green? Ideally a PR should always have green CI before anyone has a look.

Thanks!

stephenwilliams22 and others added 2 commits October 2, 2026 12:01
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@stephenwilliams22

Copy link
Copy Markdown
Collaborator Author

@stephenwilliams22 could you please always ensure that the pre-commit checks are green? Ideally a PR should always have green CI before anyone has a look.

Thanks!

@Zethson sorry about that. we should be good to go now.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants