Carry one reader's progress across a library split: export and import - #45
Merged
Merged
Conversation
Deploying codex with
|
| Latest commit: |
fedb512
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://5a4d1573.codex-asm.pages.dev |
| Branch Preview URL: | https://feat-reading-progress-transf.codex-asm.pages.dev |
…y split Every piece of per-user reading state is keyed on books.id / series.id, and both are minted fresh whenever a library is rescanned under a new root, so splitting a library loses progress, completions, sessions and ratings. GET /api/v1/reading-progress/export serialises the four tables against keys that survive a re-root: external ids, the series path, a series-relative book path, file name and hashes. Sessions are included by default because they are the only source of reading statistics; orphaned rows are skipped because nothing in the file could say which book they belonged to. POST /api/v1/reading-progress/import resolves those keys back through the same sharing-tag and access-group visibility filter as an ordinary read, so a book the importer cannot see is reported unmatched instead of leaking through a permission error. Matching never guesses between candidates and reports stem matches separately. Writes run one transaction per series under an explicit conflict policy, reuse exported completion and session ids so a second import is a no-op, and adopt rows orphaned by a hard delete instead of duplicating them. A dry run walks the same decisions without opening a transaction, and an unknown format or newer version is rejected with a 400.
Gives users the export/import endpoints without hand-crafting requests. The page downloads an export, and for import it parses the chosen file, runs a dry run, and only unlocks "Apply import" once a preview has been shown for that exact file and option set: changing any option relocks it, because the preview no longer describes what would be written. The report table shows every series and book disposition so ambiguous, unmatched and stem matches are visible before anything lands. The file is read with FileReader rather than Blob.text(), which the jsdom used by the test suite does not implement.
Covers the export, reorganise, dry-run, apply workflow end to end, every import flag and its default, how matching resolves series and books, and the rule that the export must be taken before the old library is deleted: orphaned history survives the delete, but only the export can say which book it belonged to.
… import Splitting a library usually moves the files first, and the next scan of the old library marks every moved book deleted without purging it. That is exactly when a reader takes the export, but the export looked books up through a helper that hides soft-deleted rows, so it carried nothing for the books it exists to move. It now loads them directly, in batches. Ratings were exported without a timestamp, so the default "newest" policy could not compare them and always overwrote: re-importing an old export silently undid a rating changed since. The document now carries rating_updated_at, newest and furthest replace only a strictly older rating, a file without the timestamp never overwrites, and the imported time is kept so a second import of the same file is a no-op. The PostgreSQL scenarios share one database, so the idempotency and visibility checks now compare row counts against a baseline instead of absolute numbers. With reattach real on this branch, the purge dialog, its endpoint description and the user guide point at importing an export as the alternative to deleting removed-book history.
…d input Ratings in an import file bypassed the 1..=100 range the rating endpoint enforces. Ratings feed an average every user sees, so one crafted file could skew it for everyone, and a large value overflowed the sum behind it. Out-of-range ratings are now a 400 naming the series. Session durations are clamped to their span as the normal write path does. Two entries landing on one book (a file moved inside its series appears under its old and new path) were each planned against the database as it stood, so both planned an insert, the second hit the unique index and the whole series rolled back, while the dry run had promised both. The import now keeps the state it has already planned and decides each entry against it under the conflict policy, rolled back with a series that fails, so the preview and the real run agree. On a same-instance split the old library's series, every book soft-deleted, still matched by path and stopped the search with nothing to write onto. Series with nothing on disk are no longer candidates, and history still on a soft-deleted book is moved to the matched book rather than skipped. Also: the stem step no longer checks hashes, which a repack always changes; planning failures are reported per series instead of aborting an import that has already committed earlier series; lookups are batched per series instead of one query per row; the import route accepts 64 MB instead of axum's 2 MB; series paths are exported with forward slashes; a path outside its library exports as the file name rather than an absolute server path. The settings page locks options while a request is in flight, relocks Apply after a failed apply, allows imports ten minutes, explains a 413, and keys report rows uniquely.
AshDevFr
force-pushed
the
feat/reading-progress-transfer
branch
from
September 27, 2026 22:19
13f4691 to
fedb512
Compare
AshDevFr
changed the base branch from
feat/reading-history-survives-book-delete
to
main
September 27, 2026 22:19
API contract changesCompared against Report |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Builds on #44 (merged): its nullable
book_idis what lets orphaned history survive a hard delete, andthis PR's reattach path depends on it.
Every piece of per-user reading state is keyed on
books.id/series.id, and both are minted freshwhen a library is rescanned under a new root. Splitting
/manga/shonen/Narutointo/shonen/Narutotherefore loses everyone's progress. This adds a self-service, file-based way to carry one reader's
own state across the move.
Export:
GET /api/v1/reading-progress/exportOne JSON document with the caller's
read_progress,read_completions,reading_sessionsandseries ratings, keyed by what survives a re-root rather than by database id: every external id,
the series path relative to its library, and each book's path relative to its series folder, file
name and hashes. Completions and sessions keep their UUIDs, which is what makes re-import
idempotent.
include_sessionsdefaults to on, because sessions are the only source of everyreading statistic.
Soft-deleted books are exported. In a real split the files move first, and the next scan of the
old library soft-deletes every moved book. That is exactly when a reader takes the export, so hiding
them would export nothing for the books the file exists to carry.
Measured on 5,000 books with progress, a completion and a rating each: 3.85 MB with one session
per book, 5.56 MB with two. Sessions dominate.
Import:
POST /api/v1/reading-progress/importSeries matching. Each series is matched by external id (in the request's
source_preferenceorder), then by path, then by normalized name.
Book matching. Within the series, books are matched by series-relative path, then file name,
then stem. Stem catches a
.cbrrepacked to.cbz, is reported separately, and is applied onlywith
accept_stem_matches.hash_modecontrols hashes:offignores them.verifyrejects a path or name match whose hash differs.matchalso finds books by hash.Nothing is guessed. More than one candidate at any step is
ambiguous, and nothing is writtenfor it.
Visibility. Every series candidate goes through
ContentFilter::for_user, the same predicateordinary reads use, before candidates are counted. A series or book the importer cannot see
resolves as
unmatched, never as a permission error, and nothing is written against it. It alsocannot turn a real match into
ambiguous. An integration test covers this.Conflicts.
conflict_policyisnewest(the default),furthest,skip_existingoroverwrite, and applies to progress and ratings. Ratings now carryrating_updated_at, sonewestcan tell a stale rating from a fresh one. A file without that timestamp never overwritesan existing rating except under
overwrite.Reattach. When a completion or session in the file already exists as the importer's own row
but is not on a live book, it is moved onto the matched book instead of being skipped. "Not on a
live book" means either:
book_idbecause of Keep reading history when its book is hard-deleted #44; orThat is what makes both orders safe on one instance: import after deleting the old library, or
before.
Transactions and dry run. Each series is applied in its own transaction, and the report says
which series committed. A dry run runs the same decisions, never opens a transaction, and returns
the identical shape.
Colliding entries. Two entries in the file can land on one book, for example a file that moved
inside its series appears under both paths. The import keeps the state it has already planned and
decides the second entry against the first under the conflict policy. The preview and the real run
therefore agree, and the entries never collide on the unique index.
Values the normal write paths reject. Ratings outside 1..=100 are a 400 naming the series.
Ratings feed an average every user sees, so this is the one place an import could otherwise affect
other users. Session durations are clamped to their span, as
appenddoes.Limits. The route accepts 64 MB instead of axum's 2 MB, which stops at roughly 2,500 books.
The client allows ten minutes for the request.
Settings → Reading Progress
while a request is in flight, and Apply relocks after a failed apply, which may have committed
part of the file.
The user guide covers the split workflow end to end, including the export-before-delete rule.
The purge dialog, its endpoint description and the stats guide from #44 now point at importing an
export as the alternative to deleting removed-book history.
Review notes
The first version of this branch was written by a delegated agent. I rebased it onto #44, reviewed
it, and ran an independent review, which found no visibility leak but did find real defects. All
are fixed here, each with a regression test:
Worth your judgement:
the API. The plan specified it that way, and it reads naturally as a portable interchange format.
If you'd rather keep the API uniformly camelCase, that's a mechanical change, but it's cheapest
before this ships.
passvalues are merged into the target book's log as exported. In the common case the newbook has no sessions, so this is exact. If the new book already has reading from a different
pass, the next write refolds over the combined log. Nothing is lost, but
read_progressand thelog can disagree until then.
Verification
make test-fast: 5464 passedmake test-fast-postgres-run: 5505 passed, includingreading_progress_transfer_postgres(split round trip, idempotency, visibility, colliding entries, same-instance split)
web:npm run test:run3715 passed,npm run lint,npm run buildcleancargo clippy --workspace --all-targets -- -D warningsclean