Skip to content

✨ add session sampling and error-session capture (sessionOnError) - #29

Open
Fiona2016 wants to merge 14 commits into
publishfrom
feat/session-on-error
Open

Fiona2016 wants to merge 14 commits into
publishfrom
feat/session-on-error

Conversation

@Fiona2016

Copy link
Copy Markdown
Collaborator

Summary

Adds session sampling and error-session capture to the Electron SDK.

  • sessionSampleRate (0–100, default 100): sessions are now drawn in the main process. The default keeps the existing behaviour, where every session is collected.
  • sessionOnError (default false): sessions the draw did not pick are still collected, but held in memory in the main process. They are uploaded only if the session reports an error: the last 60 seconds of main-process and renderer events are released once after a 0–3 s per-session delay, then collection continues normally. Sessions that never error are discarded.
    • Budget: 60 s window, 64 KiB / 200 non-view events, 50 views.
    • Eviction order: successful requests and long tasks first, errors last.
    • Release order: views oldest first, then errors, then the rest.
    • Only non-SDK errors of the current session trigger a release. Renderer errors dropped by the renderer's beforeSend never arrive.
    • Release happens immediately on before-quit and on a main-process uncaught exception.
    • A native crash in a held session is still reported on the next launch, together with its last view.
  • Events of held sessions report _dd.configuration.session_sample_rate: 0, and their views carry session.sampled_for_error: true.
  • For sessions that were not sampled, the bridge reports no session to renderers, so renderers do not record or upload replay for them.

Known limitation

Held sessions still upload renderer Session Replay when sessionReplayDirectUpload is enabled, because direct upload bypasses the main process. The README advises against combining it with sessionOnError.

Testing

  • Unit tests: 783 passed (vitest). Typecheck, rollup build and prettier pass.
  • The only remaining ESLint error is pre-existing, in an untouched file.
  • Playwright end-to-end: 61/61 pass, including 10 new scenarios with negative controls:
    • no error → 0 events;
    • main-process error → buffered main and renderer events released with markers;
    • renderer beforeSend dropping the error → nothing;
    • session end without an error → discarded;
    • uncaught exception → immediate release;
    • rate 0 without the switch → nothing;
    • native crash in a held session.

Add `sessionSampleRate` (0-100, default 100; an out-of-range value fails
init) and `sessionOnError` (default false). Each session is drawn once in
the main process and keeps its decision when resumed: drawn sessions report
as before, sessions not drawn report nothing, and with `sessionOnError` the
sessions the rate missed are held in memory and uploaded only once they
report an error.

- SessionManager draws the tracking type and persists it, with the error
  mark, in `_dd_s` and in the session history.
- SessionContext discards the events of sessions not drawn and the late
  events of withheld sessions that ended without an error, and stamps
  `session.sampled_for_error` on views and
  `_dd.configuration.session_sample_rate` (0 for on-error sessions) on
  main-process and renderer events alike.
- WithheldEventBuffer sits between assembly and the RUM batch: last 60s,
  64 KiB / 200 events plus 50 views, tiered eviction, release ordered
  views / errors / rest behind a per-session jitter of up to 3s, frozen
  window once scheduled, oversized releasing errors sent on their own,
  discard on session end, immediate release on session end, quit and
  uncaught main-process exceptions. Telemetry bypasses it.
- A native crash reported on the next launch releases its withheld
  session and brings its view, rebuilt from the view history.
- The bridge answers no session id for sessions not drawn, so renderers
  record nothing for them.
- Telemetry gains a debug log, used to report each release.
Scenarios against the fake intake: nothing uploaded for a withheld session
that never errors; release of main-process and renderer history with the
on-error markers; renderer errors dropped by beforeSend do not release;
discard on session end; release at once on an uncaught exception; native
crash of a withheld session reported with its view; and the negative
controls (drawn session unmarked at its rate, undrawn session uploads
nothing and renderers see no session).
- SessionManager writes its state through one promise chain, so a slow
  activity write can no longer land over the error mark; the current
  session's draw and mark live in a single copy.
- WithheldEventBuffer loses its unused stop(); an oversized error is
  measured once.
- RumCollection hands the ViewContext to ViewCollection and to the crash
  path directly, sharing its MainView type.
Accept extra released event types such as a renderer long task, and give
the scenarios, which sleep past the release jitter, a longer budget.
…release before returning

- Mark a withheld session's error on the history entry in force at the time
  of the error, plus the active entry when the session is the current one,
  rather than on every entry of the session. An error of the current launch
  no longer makes a native crash of the previous launch look already
  released, which left the crash without its view and dropped it at release.
- Chain the session file delete behind the queued state writes, so a save
  queued just before the session ended cannot recreate the file and resume
  an ended session on the next launch.
- Write the release triggered by an uncaught exception or a quit before
  returning (BatchProducer.postSync): the queued write never ran when the
  host exited from its own uncaughtException listener.
- Never evict the view that holds an error past the view limit: the error
  was dropped at release for want of its view.
- Persist the rate a session was drawn at and report that one, so a session
  resumed under a changed configuration, or saved before sessions were
  sampled, is not extrapolated by a rate that did not draw it.
- Attribute the telemetry of a session the draw did not keep to no session,
  as the browser SDK does.
- Reword the sessionOnError warning, and the README on what an exit does.
…on only once it is released

- Move exit-time durability from the withheld buffer to the transport: the
  batch producer keeps what is posted in a queue, and on APP_MAY_EXIT (an
  uncaught exception, or now a quit) Transport writes the queue before
  returning. The buffer's own synchronous sink is gone with it. An expiry
  release queued just before a quit, an oversize error forwarded on its own,
  and the fatal error of a drawn session are no longer lost to the exit.
- Name batch files strictly increasingly: several rotations within one
  millisecond shared a name, and the later .log replaced the earlier one.
- Mark the session as errored when its buffer is released, not when the
  error arrives: the mark is what tells a crash reported on the next launch
  that the crashed view reached the batch, and before the release it had
  not, so the crash was reported without its view.
- Never evict the view just updated past the view limit: the view a crash
  brings along arrives before the crash, and was evicted by the time the
  crash needed it.
- Write the session file only while the session is active, and not from an
  activity read that completed after it expired: either write brought an
  ended session back on the next launch.
- Close the previous launch's history entry on startup, resumed session or
  not: an open entry is never pruned, and the file grew with every launch
  that resumed its session.
…ts view

A detail whose view was no longer held — evicted past the view limit, or a
renderer event dated in the previous session while the renderer had not
yet noticed the renewal — was dropped at release. An error session must
not lose its error or its history; the views only order the release. An
ended view with no detail left is still pruned.
… is issued

- An event left the batch queue before its append was issued, while the
  directory check or a rotation was still awaited; an exit flush in that
  window found nothing to write and lost it. It now leaves the queue only
  as the append is issued, and the queue skips it if the exit flush wrote
  it first.
- Orphan recovery never renames a .tmp onto a .log that exists: an append
  still in flight could recreate a .tmp a synchronous rotation had just
  renamed, and the next launch replaced the earlier batch with it.
- For a crash of a session that was resumed and is current, leave the mark
  to the release of what this launch holds, like any other error: a mark
  written while the rebuilt view still waited in the buffer told a second
  crash that the view had reached the batch when it had not.
- Run upload cycles one after another instead of skipping a flush while a
  cycle runs, so that flush() resolves only once what it was called for is
  uploaded; the e2e scenario no longer pads it with a sleep.
… and never share a file between the two writers

- An event that cannot be serialized (a BigInt in an error context, say)
  rejected the batch queue's chain for good, and a failed upload cycle did
  the same to the upload chain: nothing was written or uploaded after it.
  The event is dropped like a failed write, and a failed cycle is reported
  to the flush that ran it while the next cycle runs regardless.
- A large event is appended in several chunks; a line the exit flush wrote
  to the same file between two of them corrupted both. The exit flush now
  starts a new batch while an append is in flight and leaves that file to
  the append, which rotates it once done.
- A size rotation that completed after the exit flush had rotated the same
  batch and started a new one reset the batch state, leaving the new batch
  an untracked .tmp until the next launch. The reset now applies only when
  the batch rotated is still the current one.
- Telemetry formatted an error before checking whether it would send it,
  so a sampled-out telemetry could throw on a value it cannot serialize
  from inside an error path and lose the error being reported.
- The main process's sampling marker is assigned on bridged events rather
  than merged, so one a renderer supplied is replaced, absent included.
…its uncaughtException listener

The existing uncaught-exception scenario keeps the process alive and
flushes, so it passed with the exit-time write removed. The new scenario
ends the process from the application's own listener, registered after
the SDK's, and reads the batch files off disk: the held history and the
fatal error are there, and nothing reached the intake.
…ith the exit flush

- The withheld buffer measured events with browser-core's byte count,
  which takes an ASCII fast path and otherwise reaches for
  window.TextEncoder — undefined in the main process. Holding an error
  whose text was not ASCII threw, the session was never released, and an
  exit then had nothing to write. Bytes are counted with Buffer.
- The exit flush wrote the released batch synchronously but left the
  session's error mark to an asynchronous write, so a host that ended the
  process from its uncaughtException listener left the next launch
  resuming the session as withheld again. The flush now writes the session
  state and history too, and the history is serialized when a write runs
  rather than when it is queued, so an earlier queued write never lands an
  older state.
- The view of an error forwarded on its own for exceeding the budget had
  no protection against eviction past the view limit; a renderer view
  arriving during the jitter could push it out before the release.
- The synchronous batch rotation now also avoids a .log that exists, as a
  clock set back can make a new launch reuse an earlier launch's name.
- The main view counts only its own events: a crash reported on the next
  launch is rebuilt with its own view and count, and was counted twice.
- Docs: an append already issued is not completed by the exit flush; the
  application's uncaughtException listener must be registered after init.
- Tests: the host-exit scenario asserts the child's exit code; mock
  implementations no longer leak across cases; capped telemetry case.
…s file, and age the buffer by a monotonic clock

- The exit flush wrote the session state and history in place, where a
  write of the same file already issued could overwrite or interleave
  with it. Both are now written to a file of their own and renamed into
  place.
- A session ended with stopSession() right before a fatal exit kept its
  file: its delete was queued behind the writes. The exit flush deletes it.
- The view counter guard read the event's view before its type; a
  telemetry event assembled after the main view closed has none, and
  reporting an SDK failure threw. The type decides first.
- The withheld buffer aged its events by the wall clock, so a clock
  correction before the error pruned the minute it promises. It ages by
  the monotonic clock, as the browser SDK does.
… of its own

A native crash is reported on the next launch and attributed to the view
in force when it happened; when that view was no longer in the view
history, the crash was discarded at assembly for want of a container, and
a withheld session was left unreleased on top. A session must not lose
its crash: the session is released regardless, and the crash goes out
under a view id of its own, with the session's markers and a rate of 0,
when none can be rebuilt. The RUM hooks take the view a raw event names
itself as a fallback to the one in force at its time.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant