Skip to content

feat: keep sessions and replays missed by sampling until the session errors - #36

Open
Fiona2016 wants to merge 13 commits into
publishfrom
feat/session-on-error
Open

Fiona2016 wants to merge 13 commits into
publishfrom
feat/session-on-error

Conversation

@Fiona2016

@Fiona2016 Fiona2016 commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Adds error-session capture to RUM and Session Replay.

RumConfiguration.Builder.setSessionOnError(Boolean)

Default false. Remote key: rum.sessionOnError.

  • What it does: sessions the session sample rate does not pick enter a new WITHHELD state. Their events are assembled and run through the event mappers, then held in memory by WithheldEventWriter instead of being written.
  • On error: the last 60 seconds of events are released once, after a 0–3 s delay derived from a hash of the session id. Views go first (oldest first), then errors, then the rest. Every held event is released; views only set the order. The session is then collected normally.
  • Without an error: the session is discarded when it expires, renews or stops. Late events of discarded sessions are dropped.
  • Limits: 60 s window, 64 KiB / 200 non-view events, 50 views. Successful requests and long tasks are evicted first, errors last.
  • Release triggers:
    • Only errors that survive the mappers trigger a release.
    • A JVM crash releases synchronously inside the crash write, including when the crash event is oversized.
    • setForcedSession() releases immediately.
    • A pending release is sent as soon as the app goes to the background.
  • Session listener: reports these sessions as discarded, and getCurrentSessionId returns null until they are released.
  • NDK crashes and fatal ANRs: reported on the next launch for these sessions, at rate 0 with the marker.

setSessionReplayOnError(Boolean) on the Session Replay configuration

Default false. Remote key: rum.sessionReplayOnError.

  • What it does: replays the replay sample rate does not pick are recorded, but records and images are held in memory per session. Replays of sessions whose events are withheld are held too.
  • Window: the held replay is cut only at full snapshots and keeps the last 60 s across views, capped at 4 MiB. The first released segment always starts with Meta, Focus and a full snapshot.
  • Release: the held replay is released together with the withheld events, after their delay. A replay never reaches the intake before the session's views.
  • Replay counts: records_count and has_replay count only records that were actually written. On release, session.has_replay is set on released views and events whose view has replay records.

Reporting

  • Events of withheld sessions report _dd.configuration.session_sample_rate: 0.
  • View events carry session.sampled_for_error and session.sampled_for_error_replay. session.sampled_for_replay is reported only when an on-error mode applies to that session.
  • Customers who use neither switch produce byte-for-byte unchanged events.

Remote configuration

  • A rate of 0 with the switch on no longer resets a withheld session.
  • Toggling the switch while the rate is 0 resets the session.
  • beforeSampling returning 0 also turns sessionOnError off for that session.

Behaviour change for all users

  • RumViewEventFilter still keeps only the latest version of each view in a batch, but now places it at the view's first position instead of the position of its last write. A release followed by view updates in the same batch therefore keeps views ahead of errors and details. The set of uploaded events is unchanged.

Known limitations

  • After a JVM crash, a held replay is lost with the process; the held events are released.
  • An app that is force-stopped within about a second of an error can lose the release, because the write is still in flight.

Testing

  • RUM: 2385 unit tests, 0 failures.
  • Session Replay: 2492 unit tests, 19 failures. The same 19 mapper test classes also fail on publish with JDK 21; they are a pre-existing Typeface.DEFAULT reflection error.
  • ktlint: clean on the touched files.
  • Emulator end-to-end, against a local capture intake, with each scenario paired with a negative control:
    • no error → 0 requests;
    • an error → views first, the markers set, and the rate reported as 0;
    • the mapper dropping the error → nothing is sent;
    • a JVM crash → handled;
    • setForcedSession() → released immediately;
    • remote configuration at rate 0 with the switch on → the session is not reset;
    • replay type 3 → released;
    • replay type 5 → released after the events, with has_replay set;
    • a cut at 75 s → the first segment starts with Meta, Focus and a full snapshot;
    • a normal Session Replay customer → the view session keys are unchanged.

…n error

Adds RumConfiguration.Builder.setSessionOnError and the remote
configuration keys rum.sessionOnError / rum.sessionReplayOnError
(absent keeps the init value; the console wins over init).

A session the session sample rate does not keep is drawn as WITHHELD
when the switch is on: it is assembled as usual but its events go to an
in-memory WithheldEventWriter instead of the batch. Events are held
after the event mappers have run and serialized, so an error dropped by
a mapper releases nothing. The buffer keeps the last 60s, at most 64KiB
and 200 non-view events plus 50 views (latest update per view, current
view never evicted), evicting successful requests and long tasks first
and errors last, newest first.

The first error of the session releases the buffer after a per-session
jitter of 0-3s (String.hashCode based), with the window frozen when the
release is scheduled; views go first oldest to newest, then errors, then
the rest. A JVM crash releases synchronously within the crash write, an
oversized error is written on its own, and setForcedSession releases
immediately. When the session expires, is renewed or stopped, an
unreleased buffer is dropped and its id remembered (last four) so late
events of that session are dropped too; a session that had errored is
released instead. Withdrawn consent drops what is held.

Withheld sessions report a session sample rate of 0, view events carry
session.sampled_for_error, the session listener sees them as discarded
and getCurrentSessionId answers null until release. The last view is
still written locally, so a native crash or fatal ANR is reported at the
next launch with the marker and rate 0.

A console reset no longer ends a session kept on error while the rate
is 0 and the switch is on; a switch change at rate 0 now resets the
running session, and beforeSampling returning 0 turns the switch off.
… session errors

Adds RumConfiguration.Builder.setSessionReplayOnError (remote key
rum.sessionReplayOnError, latched per session at the draw).

RUM now tells Session Replay, on the session bus message, whether the
session's events are kept on error, the replay switch it was drawn
under, and whether it has reported its error. Session Replay keeps its
own replay draw and holds the records of two kinds of session in memory
instead of writing them: a collected session whose replay the rate
missed while the switch is on, and a session whose events are withheld,
whichever way the replay draw went, since its replay has nothing to
attach to until the events are released. The error releasing a
replay-only session is judged after the event mappers, like the events'
own; forcing the session releases too.

Held records are cut only at full snapshots, keeping from the newest full
snapshot at least a minute old, so the release spans the last minute
across views and starts playable. View record counts and has_replay are
only updated for records actually written, so nothing held and thrown
away is ever claimed. A session that ends unreleased drops its records,
and records of that session still queued are dropped too.

View events carry session.sampled_for_error_replay for replays kept on
error and session.sampled_for_replay, which counts a held replay whose
events are held too.
Session Replay resources now go through the record writer, which holds
them with the records of a withheld session and writes or drops them
together; a resource hash is remembered as sent only once it is written.

The held replay is released when the session's events actually are -
after the jitter, by an explicit bus message from the event buffer - not
when the error is first seen. Released events and the releasing error
claim session.has_replay when their view has replay records (sent or
held), counted by Session Replay apart from records_count. A release
waiting for the jitter goes at once when the app leaves the foreground,
and an oversized crash releases the history right after it.

The release forwards every held detail, views only setting the order,
so an error whose view was never held or was evicted still goes out.
The held replay gains a 4MiB bound (images bounded apart), cut at full
snapshots, and a release cut at a periodic full snapshot starts with
the view's meta and focus records.

sampled_for_replay is only reported for sessions in an on-error mode,
and Session Replay leaves its feature context untouched when nothing is
withheld, so views of customers who did not opt in are unchanged.
…n in a batch

The upload filter keeps only the latest version of each view in a batch.
It used to keep that version where it was written, so when a released
error session's buffer (views first, oldest first, then errors and
details) shared a batch with the live updates written right after the
release - always the case for a crash, which updates the view's crash
count - the views ended up last and out of start order. The intake
builds the session out of the first view it sees.

The latest version now takes the place of the view's first occurrence
in the batch; which versions survive is unchanged.
…ur consent and never evict a crash

Session Replay now keeps aside the replay of a session whose events were
withheld until RUM says whether that session was released or thrown away,
instead of discarding it as soon as the next session announces itself:
RUM's word lands on the storage thread and can arrive after the next
session has started. RUM sends an explicit discard for such a session, and
a stopped session draining beside its successor no longer announces itself
as current.

Images captured while a replay is withheld are held across sessions and
sent only with released records that show them, since the recorder
captures an image once per process. Records and images captured while
tracking consent is withdrawn are neither held nor sent, and what was held
is dropped. The replay byte budget is enforced on every record.

A crash is never evicted from the withheld event buffer, even when it does
not fit alongside an earlier large error. A session kept on error is only
ended by the emergency stop, not by a console rate leaving zero.
…rop buffers on consent withdrawal

Session Replay keeps aside every replay whose session ended while held,
including a collected session's replay held on error, until RUM says what
became of it; RUM now reports the end of such a session too. Several ended
sessions can wait at once, since storage may lag more than one session
behind. A stopped session draining beside its successor no longer draws a
session of its own, which ended the running session's buffer, and a late
error of a stopped session no longer unmarks the running session's release.
Session Replay handles RUM's announcements and fates under one lock.

Both features now observe tracking consent directly: what was held in
memory under a consent since withdrawn is dropped even when no event comes
by to notice. A replay span over the byte budget is dropped rather than
kept, a sent record flushes the held images it shows, and a view's meta
and focus survive the consent drop so the resumed replay stays playable.
…withheld session's data

A withheld session that ends without an error now also forgets the view it
wrote locally for the native crash reporter, so a crash of a later
uncollected session is not reported against it. Both writers remember more
discarded sessions, since a stopped session keeps draining its requests
while the sessions after it come and go. A stopped session keeps its own
state for what it still drains instead of expiring. An SDK stop flushes a
release still waiting for its jitter.

Session Replay notes RUM's word on an ended session the moment it is
given, so a released session is never the one thrown away when too many
ended sessions wait, and a view's meta and focus survive the drop of a
span over the byte budget. A view of a session whose replay was released
claims the replay as soon as the records exist, so the final view of a
stopped session does not replace the released one without it.

The web view consumers leave the replay and the replay claim alone while
the native replay of the session is withheld.
…t rate 0 on app-launch vitals

The recorder captures each image once per process. An image held for a
replay kept on error and then dropped unsent - evicted over the image
budget or cleared when consent is withdrawn - was never captured again, so
a replay released later showed a placeholder for it. The writer now tells
the recorder what it dropped; the recorder forgets those ids and resolves
such an image again from its drawable the next time it is shown, instead
of answering from its caches.

The app-launch vitals of a session kept on error report a session sample
rate of 0, like every other event of such a session.
… view's replay markers once

The core takes a feature out of its registry before stopping it, so the
flush of a release still waiting for its jitter could not look the RUM
scope up at stop time: the scope is kept from initialization.

A view of a withheld session keeps its Session Replay entry when it
completes, so it can still claim the replay held for it when the session
is released; Session Replay drops the entry once nothing is sent or held
for the view. A view's replay markers stay as resolved once they are
true, so a late update written after another session took over the
replay context still describes its own session.

The replay switch alone no longer restarts a running session on an
immediate configuration change: a replay draw is made once per session.
…eport a withheld session's last view however old

A session already released holds nothing more once consent is withdrawn:
its events go to the batch, where consent decides, instead of through the
buffer, where a release granted later would have carried out what was
collected while consent was withdrawn.

An image first captured while consent is withdrawn is forgotten by the
recorder like one dropped from the held store, so it is captured again;
the forgotten ids are bounded, and only the view in progress keeps its
start across a consent drop. More discarded sessions are remembered,
since a request can outlive many sessions.

The native crash reporter writes the last view of a session kept on error
with the crash however old the view is: that session uploaded nothing
before it crashed, so this view is the only one the intake will get.
…hing for a stopped session thrown away

The SDK stop ends the session kept on error with it: released if it had
reported its error, thrown away if not, along with the view it wrote
locally - a crash after the stop must not be reported against a session
that never errored. The write scope only queues that work, and the core
shuts its executor down without draining it once the features are stopped,
so the stop waits for the queue to get there, the way a crash does.

A stopped session still draining its requests hands its views nothing to
write with while it is withheld: it was thrown away at the stop, and
remembering its id for a while could not cover a request that outlives
dozens of later sessions. A stopped session that had errored is collected
by then and keeps writing.
Settling the withheld session at the stop looked the write scope up
synchronously, which waits on the context thread without a bound; a stop
asked for from the RUM thread - a session listener, say - found that
thread waiting on the RUM thread in turn, and the stop never returned,
holding the SDK registry lock with it.

The stop now queues the work through the ordinary write scope and waits,
with a bound, for the persistence thread to have run it. A stop asked
for from inside the SDK's own pipeline waits the bound out and goes on
without the work, which is what it did before, minus the hang.
…sion Replay stops

The features stop in no particular order. Session Replay only heard that
a replay may go out when RUM actually released the events, after the
jitter, and a stop in between left it stopping with the replay still held
and nobody left to tell it. RUM now says so the moment the releasing error
is written, Session Replay notes it, and its own stop writes every replay
whose session reported an error, throws the rest away, and waits - with a
bound - for that to have run.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant