Threadline emits :telemetry events for every capture, health, export, retention, and operator-surface milestone it drives internally. No event carries row values, actor identifiers, correlation ids, or free-text reasons — measurements and metadata stay limited to counts, durations, outcome atoms, and structural facts such as which keys a scope map used. Attach your own handlers to build metrics, alerts, and dashboards on top of this contract.

Events

EventMeasurementsMetadataWhen emitted
[:threadline, :transaction, :committed]table_count—an AuditTransaction is committed
[:threadline, :action, :recorded]status—Threadline.record_action/2 completes, whether it succeeds or fails
[:threadline, :health, :checked]covered, expected_uncovered, uncovered—Threadline.Health.trigger_coverage/1 returns
[:threadline, :health, :checked, :error]—exceptiona polled coverage check raises
[:threadline, :health, :findings_checked]errors, warnings—Threadline.Health.trigger_findings/1 returns
[:threadline, :operator_surface, :authorize]resultpath, scope_keysan operator-surface mount or request is authorized, denied, or errors
[:threadline, :operator_surface, :export_authorize]count, result—an export-specific authorization check raises
[:threadline, :operator_surface, :actor_ref_mismatch]count—the session actor and the scope-derived actor disagree
[:threadline, :export, :completed]duration, row_countformat, truncatedan export (eager CSV/JSON, the async orchestrator job, or the chunked operator-surface download) finishes successfully
[:threadline, :export, :failed]duration, row_countformat, error_kind, exceptionan export (eager CSV/JSON, the async orchestrator job, or the chunked operator-surface download) fails
[:threadline, :retention, :purge, :start]monotonic_time, system_timedry_run, telemetry_span_contextafter purge/1's input checks pass, when the purge work begins
[:threadline, :retention, :purge, :stop]batches_run, deleted_changes, deleted_transactions, duration, monotonic_timedry_run, telemetry_span_contextwhen the run or preview returns
[:threadline, :retention, :purge, :exception]duration, monotonic_timedry_run, kind, reason, stacktrace, telemetry_span_contextwhen the database raises mid-run
[:threadline, :retention, :batch_purged]deleted_changes, deleted_transactions, duration—after a purge_loop step's change delete_all and full orphan drain both return, once per step including the terminating empty one

Attaching handlers

Attach one handler per family (or :telemetry.attach_many/4 across the whole family) with a module-function capture, not an anonymous function — anonymous functions cannot be hot-code upgraded and are harder to detach individually in tests.

Capture — transaction commits and recorded actions:

:telemetry.attach_many(
  "my-app-capture",
  [
    [:threadline, :transaction, :committed],
    [:threadline, :action, :recorded]
  ],
  &MyApp.Instrumentation.handle_capture_event/4,
  nil
)

Health:

:telemetry.attach_many(
  "my-app-health",
  [
    [:threadline, :health, :checked],
    [:threadline, :health, :checked, :error],
    [:threadline, :health, :findings_checked]
  ],
  &MyApp.Instrumentation.handle_health_event/4,
  nil
)

Export — [:threadline, :export, :completed] and [:threadline, :export, :failed] fire once per logical export, from Threadline.Export.to_csv_iodata/2, to_json_document/2, the async Threadline.Export.Orchestrator job, and the chunked operator-surface download. If your host builds its own export flow on top of the lower-level stream primitives instead of these four entry points, your code owns emitting its own completion/failure events for that unit of work — Threadline does not emit on your behalf there.

:telemetry.attach_many(
  "my-app-export",
  [
    [:threadline, :export, :completed],
    [:threadline, :export, :failed]
  ],
  &MyApp.Instrumentation.handle_export_event/4,
  nil
)

Retention — the purge span (:start / :stop / :exception) plus the per-batch event:

:telemetry.attach_many(
  "my-app-retention",
  [
    [:threadline, :retention, :purge, :start],
    [:threadline, :retention, :purge, :stop],
    [:threadline, :retention, :purge, :exception],
    [:threadline, :retention, :batch_purged]
  ],
  &MyApp.Instrumentation.handle_retention_event/4,
  nil
)

A disabled or misconfigured Threadline.Retention.purge/1 call emits none of these — the span opens only after policy resolution and the disabled check both pass, so there is nothing to subscribe to until a purge actually runs. [:threadline, :retention, :purge, :exception] means PostgreSQL raised mid-run; its reason/stacktrace metadata is forwarded for incident diagnosis only — see Keep row data out of your handlers before logging it anywhere durable. Each [:threadline, :retention, :batch_purged] event fires before the outer caller's own transaction (if any) commits, so a handler that reads its own database inside the handler can observe a batch that is not yet externally visible. A dry run's [:threadline, :retention, :purge, :stop] counts are a preview estimate, not a result — for the row totals an actually-completed purge deleted, sum deleted_changes/deleted_transactions from the batch_purged events instead.

Operator surface — authorization outcomes and the actor-mismatch counter:

:telemetry.attach_many(
  "my-app-operator-surface",
  [
    [:threadline, :operator_surface, :authorize],
    [:threadline, :operator_surface, :export_authorize],
    [:threadline, :operator_surface, :actor_ref_mismatch]
  ],
  &MyApp.Instrumentation.handle_operator_surface_event/4,
  nil
)

[:threadline, :operator_surface, :actor_ref_mismatch] is a pure incidence counter (%{count: 1}, no metadata) — it tells you a mismatch happened, not which actor or scope. result on authorize/export_authorize and status on [:threadline, :action, :recorded] are atom-valued measurements (:granted / :denied / :error, or :ok / :error), not strings — match on the atom, not on a string you format yourself.

Metrics with telemetry_metrics

The example below is adopter code — Threadline never adds the telemetry_metrics dependency itself. Durations (duration, monotonic_time) are native time units; declare unit: {:native, :millisecond} so your metrics backend renders milliseconds instead of raw VM-native units.

[
  Telemetry.Metrics.summary("threadline.export.completed.duration",
    event_name: [:threadline, :export, :completed],
    measurement: :duration,
    unit: {:native, :millisecond}
  ),
  Telemetry.Metrics.counter("threadline.export.failed.count",
    event_name: [:threadline, :export, :failed],
    tags: [:error_kind]
  ),
  Telemetry.Metrics.summary("threadline.retention.purge.duration",
    event_name: [:threadline, :retention, :purge, :stop],
    measurement: :duration,
    unit: {:native, :millisecond},
    tags: [:dry_run]
  ),
  Telemetry.Metrics.counter("threadline.retention.batch_purged.count",
    event_name: [:threadline, :retention, :batch_purged]
  )
]

Handlers must not raise

:telemetry isolates one failing handler from the event it was attached to, not from your application's stability guarantees — a raising handler is caught, logged, and detached from every event in the attach_many call that registered it, not only the one that raised. [:telemetry, :handler, :failure] fires when this happens, naming the failed handler's id. Attach a separate handler to that event if you want to alert when one of your own handlers goes dark:

:telemetry.attach(
  "my-app-handler-failure-alert",
  [:telemetry, :handler, :failure],
  &MyApp.Instrumentation.handle_handler_failure/4,
  nil
)

A raising handler never changes Threadline's own result: export, purge, and every other Threadline operation returns its normal value regardless of whether your handler crashed while observing it.

Cardinality

Never tag a metric on an id-shaped value — table_pk, a correlation id, an actor identifier, or any other high-cardinality value turns a bounded set of time series into an unbounded one. scope_keys on [:threadline, :operator_surface, :authorize] is the keys of your host-returned scope map, not its values; if your host happens to use identity-shaped values as map keys (unusual, but not forbidden), tagging on scope_keys would reintroduce the same cardinality problem through the back door. Tag on format, error_kind, dry_run, or other low-cardinality enums instead.

path on [:threadline, :operator_surface, :authorize] is the mount's own compile-time route template (e.g. "/audit/theme") — not the live request path. If your router nests the Threadline mount under a dynamic segment (for example threadline_operator_surface("/accounts/:account_id/audit")), path still carries the un-substituted ":account_id" placeholder, never a real account id, so it is always safe to tag on.

Keep row data out of your handlers

Threadline's own events never carry row values, actor identifiers, correlation ids, or free-text reasons — that boundary is enforced in this module and proven by a dedicated property test. A handler you write can still reopen that leak: never log an event's :params (repo-query telemetry carries plaintext bind values — see Observing Threadline's queries below) or a purge-exception's reason/stacktrace verbatim to a sink your audit boundary does not cover, and never attach row data of your own onto a Threadline event's measurements or metadata before forwarding it downstream.

Observing Threadline's queries

Threadline adds no query telemetry event of its own — every SQL statement it runs goes through your own repo, so it already shows up on your host repo's [:my_app, :repo, :query] event. Attach there and match metadata.source in ~w(audit_changes audit_transactions audit_actions) to isolate Threadline's own queries from the rest of your application's:

:telemetry.attach(
  "my-app-threadline-query-time",
  [:my_app, :repo, :query],
  fn _event, measurements, metadata, _config ->
    if metadata.source in ~w(audit_changes audit_transactions audit_actions) do
      MyApp.Metrics.record("threadline.query_time", measurements.query_time,
        unit: {:native, :millisecond}
      )
    end
  end,
  nil
)

The following caveats are proven by a dedicated test, not just asserted here:

  • :source is the bare table name with no storage-schema prefix — even when Threadline is configured to store its tables outside the default threadline schema, metadata.source stays "audit_changes", never "other_schema.audit_changes".
  • :source is nil for raw SQL (Repo.query!/2) and for a query rooted in a subquery (some of Threadline's own counts use one to cap expensive aggregates) — queries in either shape are not attributable to a table through :source at all.
  • Capture itself — the PostgreSQL trigger function writing audit_changes/audit_transactions rows — runs entirely inside the database as part of executing your application's own statement against the audited table. It never appears as its own repo query event; only the statement your application issued against the host table does.
  • Never log metadata.params from this event. It carries the literal bind values of every query your repo runs, including the full row data of whatever your application just wrote to an audited table — exactly the leak Keep row data out of your handlers warns about, through a channel Threadline does not control.

Next steps