Skip to content

perf: investigate in-memory state heap footprint (~60x on-disk-to-heap expansion) #132

Description

@themightychris

Context

Surfaced during an incident (PR #131): a cold boot rebuilding in-memory state from the full published import OOM'd at the 1536 MiB V8 old-space ceiling. Mitigated by raising the heap to 3072 MiB + container limit to 3.5Gi, but the underlying footprint is worth a real look.

The smell

  • On-disk record data on published is only ~25 MB of TOML (uncompressed blob bytes across people/projects/updates/buzz/memberships/tag-assignments).
  • A cold boot building the in-memory representation needs >1.5 GB of V8 heap — roughly a 60× expansion.
  • The full-text store is better-sqlite3 FTS5 (apps/api/src/store/fts.ts), which is off-heap (native memory), so it does not explain the V8 heap blowup. The heap cost is the record maps + secondary indices.

Dataset shape (published)

  • people: 31,832
  • tag-assignments: 10,353
  • tags: 1,005
  • projects: 268
  • project-memberships: 822 / project-updates: 503 / project-buzz: 112

Worth investigating

  • Duplication: does each record exist in more than one in-memory structure (raw map + multiple indices + gitsheets' cached per-sheet dataTree snapshot)? See apps/api/src/store/memory/ (loader.ts, state.ts) and apps/api/src/store/sheet-indices.ts, plus the gitsheets openStore data-tree caching noted in apps/api/src/store/public.ts / memory/reload.ts.
  • Per-record overhead: 31.8k people at >40 KB heap each would be ~1.3 GB — are person objects carrying large/derived fields, retained TOML source strings, or boxed structures?
  • Index sizing: how many secondary indices, and are any of them O(records) string-keyed Maps holding full objects rather than ids?
  • gitsheets retention: confirm whether the parsed blob contents / tree objects are held alongside the typed record maps (a second full copy).

Goal

Bring the cold-boot heap footprint down meaningfully (target: comfortably under 1.5 GB for the current import, with margin to grow), or document why the current footprint is irreducible. This also de-risks prod, where the dataset only grows.

Related: PR #131 (the heap/limit bump mitigation).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions