Skip to content

Document how to retrieve TPB amplitudes, and the bitstring packing order - #55

Merged
Jim Garrison (garrison) merged 3 commits into
mainfrom
document-wavefunction-output
Oct 9, 2026
Merged

Jim Garrison (garrison) merged 3 commits into
mainfrom
document-wavefunction-output

Conversation

@garrison

@garrison Jim Garrison (garrison) commented Oct 9, 2026 •

Copy link
Copy Markdown
Member

Two things a user currently has to discover by reading the vendored C++ or by experimenting. Both came up in a user question about the Python bindings.

TPB has two wavefunction output paths, and they write different formats

tpb_diag returns no amplitudes, and there are two ways to get them:

sbd_data.dump_matrix_form_wf savename argument
Writes the full len(adet) × len(bdet) subspace, gathered onto one rank one file per b_comm position, each holding only that rank's block
Path exactly what you set f"{savename}{rank_b:06d}.bin"
Header none three size_t (adet_range, bdet_range, det_length), then that block's determinant words
Payload float64, row-major over α×β adet_range × bdet_range float64
Intended for analysis checkpoint/restart, paired with loadname

dump_matrix_form_wf is what almost every caller wants, and is already what this package's own SQD integration uses — sbd_solver.py sets it and reads the result back with a bare np.fromfile. But it appeared only as a one-line pybind docstring and a CLI flag in examples/tpb/run_sbd_diag.py; it was absent from the README's TPB section and from the tpb_diag docstring, both of which mention savename alone. Following the function signature therefore led to the per-rank restart layout, which is the wrong thing to parse for analysis.

This adds the comparison to the README, a Note: block on tpb_diag, a pointer on tpb_diag_from_files, a row in the TPB_SBD config table, and a fuller pybind docstring. It also records that the .dat/.txt text form is written with the stream's default formatting, so it carries roughly 6 significant digits rather than full float64 — restart.h sets precision(16) on std::cout, not on the file stream.

The word-packing convention was unstated

Bitstrings are packed from the right: the trailing bit_length characters become word 0, so the leading characters of a string longer than one word land in the last word rather than the first.

from_strings(["0011"], 20, 4)           # -> array([[3]], dtype=uint64)
from_strings(["1" + "0" * 69], 63, 70)  # -> array([[ 0, 64]], ...)
from_strings(["0" * 69 + "1"], 63, 70)  # -> array([[1, 0]], ...)

Worth stating explicitly because getting it backwards produces a valid-looking but wrong subspace rather than an error. It is documented on from_strings, which #53 lists in the API reference, rather than on the singular from_string, which it deliberately leaves out; the singular forms point at it.

Documentation only — no behavior change, hence an other release note.

GDB needed no changes: #46 already documents its amplitude retrieval and per-shard file layout in examples/gdb/README.md.

Merged main after #53

#53 landed first and changed where this documentation belongs, so the merge commit is followed by a commit reconciling the two. No files overlapped — #53 is docs/apidocs/*.rst only — but three things needed fixing:

tox -e docs passes with no warnings.


This pull request was drafted by Claude Opus 5 under my guidance.

Two things a user has to discover by reading upstream C++ or experimenting.

TPB has two wavefunction output paths that write different formats, and
nothing said so. `savename` -- the one named in `tpb_diag`'s signature --
goes to `SaveWavefunction`, which writes one file per b_comm position
holding only that rank's block, behind a three-`size_t` header and the
block's determinant words. `sbd_data.dump_matrix_form_wf` goes to
`SaveMatrixFormWF`, which gathers the full adet x bdet subspace onto one
rank and writes a bare row-major float64 array with no header at all, or,
for a `.dat`/`.txt` path, a text table labelling each amplitude with its
own alpha and beta bitstring.

The second is what almost every caller wants, and is what this package's
own SQD integration already uses (`sbd_solver.py` sets it and reads the
result with a bare `np.fromfile`). It appeared only as a one-line pybind
docstring and a CLI flag in `run_sbd_diag.py`, absent from the README's
TPB section and from the `tpb_diag` docstring, both of which mention
`savename` alone -- so following the signature led to parsing the wrong
layout. Document both, side by side, and note that the text form is
written with default stream formatting and so carries ~6 significant
digits rather than full float64.

`from_string` and `makestring` packed words with no stated convention.
Bitstrings are packed from the right: the trailing `bit_length` characters
become word 0, so the leading characters of a multi-word string land in
the last word, not the first. Easy to get backwards, and a wrong guess
yields a valid-looking but wrong subspace rather than an error.

GDB needed no changes here: #46 documented its amplitude retrieval and
file layout in examples/gdb/README.md.

Assisted-by: Claude Opus 5
@garrison Jim Garrison (garrison) added the documentation Improvements or additions to documentation label Oct 9, 2026
@garrison
Jim Garrison (garrison) marked this pull request as ready for review October 9, 2026 19:58
#53 curated the autodoc list and added field tables, which changes where
this branch's documentation belongs.

The packing convention moves from `from_string` to `from_strings`. #53
documents the bulk form and deliberately leaves the singular one out, so
the convention was written on a function the rendered docs never show --
and `from_strings`, the form users are steered to, had no packing
documentation of its own. `from_string` and `makestring` now point at it
instead of restating it. The examples are shown as the `uint64` ndarray
`from_strings` actually returns rather than as nested lists.

`dump_matrix_form_wf` now appears in #53's `TPB_SBD` field table as well
as the README's, so the README row defers to the format section instead of
describing the field a second time.

Also corrects a claim this branch made about `sbd_solver.py`: since #49 the
wavefunction write is skipped when SBD selects the carryover itself
(`carryover_type` 1 against an addon that accepts it), because the next
subspace then comes back in `carryover_adet`/`carryover_bdet` and the
amplitudes never reach Python. Saying it "uses" the dump without that
qualification implied the write always happens.

`tox -e docs` passes with no warnings, and the packing convention renders
on the `from_strings` entry.

Assisted-by: Claude Opus 5
@garrison Jim Garrison (garrison) added the stable backport potential Eligible for backporting label Oct 9, 2026
@garrison
Jim Garrison (garrison) merged commit 90afb07 into main Oct 9, 2026
15 checks passed
@garrison
Jim Garrison (garrison) deleted the document-wavefunction-output branch October 9, 2026 20:44
@garrison

Copy link
Copy Markdown
Member Author

Mergify (@Mergifyio) backport stable/1.7

@mergify

mergify Bot commented Oct 9, 2026

Copy link
Copy Markdown

backport stable/1.7

⚠️ Cannot use the command backport stable/1.7

Details

⚠ The product Workflow Automation needs to be activated to enable this feature.

@mergify

mergify Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

backport stable/1.7

✅ Backports have been created

Details

Jim Garrison (garrison) added a commit that referenced this pull request Oct 10, 2026
…der (#55) (#57)

* Document how to retrieve TPB amplitudes, and the bitstring packing order

Two things a user has to discover by reading upstream C++ or experimenting.

TPB has two wavefunction output paths that write different formats, and
nothing said so. `savename` -- the one named in `tpb_diag`'s signature --
goes to `SaveWavefunction`, which writes one file per b_comm position
holding only that rank's block, behind a three-`size_t` header and the
block's determinant words. `sbd_data.dump_matrix_form_wf` goes to
`SaveMatrixFormWF`, which gathers the full adet x bdet subspace onto one
rank and writes a bare row-major float64 array with no header at all, or,
for a `.dat`/`.txt` path, a text table labelling each amplitude with its
own alpha and beta bitstring.

The second is what almost every caller wants, and is what this package's
own SQD integration already uses (`sbd_solver.py` sets it and reads the
result with a bare `np.fromfile`). It appeared only as a one-line pybind
docstring and a CLI flag in `run_sbd_diag.py`, absent from the README's
TPB section and from the `tpb_diag` docstring, both of which mention
`savename` alone -- so following the signature led to parsing the wrong
layout. Document both, side by side, and note that the text form is
written with default stream formatting and so carries ~6 significant
digits rather than full float64.

`from_string` and `makestring` packed words with no stated convention.
Bitstrings are packed from the right: the trailing `bit_length` characters
become word 0, so the leading characters of a multi-word string land in
the last word, not the first. Easy to get backwards, and a wrong guess
yields a valid-looking but wrong subspace rather than an error.

GDB needed no changes here: #46 documented its amplitude retrieval and
file layout in examples/gdb/README.md.

Assisted-by: Claude Opus 5

* Fit the amplitude and packing docs to the API surface #53 defined

#53 curated the autodoc list and added field tables, which changes where
this branch's documentation belongs.

The packing convention moves from `from_string` to `from_strings`. #53
documents the bulk form and deliberately leaves the singular one out, so
the convention was written on a function the rendered docs never show --
and `from_strings`, the form users are steered to, had no packing
documentation of its own. `from_string` and `makestring` now point at it
instead of restating it. The examples are shown as the `uint64` ndarray
`from_strings` actually returns rather than as nested lists.

`dump_matrix_form_wf` now appears in #53's `TPB_SBD` field table as well
as the README's, so the README row defers to the format section instead of
describing the field a second time.

Also corrects a claim this branch made about `sbd_solver.py`: since #49 the
wavefunction write is skipped when SBD selects the carryover itself
(`carryover_type` 1 against an addon that accepts it), because the next
subspace then comes back in `carryover_adet`/`carryover_bdet` and the
amplitudes never reach Python. Saying it "uses" the dump without that
qualification implied the write always happens.

`tox -e docs` passes with no warnings, and the packing convention renders
on the `from_strings` entry.

Assisted-by: Claude Opus 5
(cherry picked from commit 90afb07)

Co-authored-by: Jim Garrison <garrison@ibm.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation stable backport potential Eligible for backporting

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant