ForgeOpsTracker (forge_ops_tracker v0.13.0)

Copy Markdown View Source

Elixir error reporting client for ForgeOps:

ForgeOpsTracker.init(dsn: "https://<api_key>@getforgeops.net/api/v1/events")
ForgeOpsTracker.install_handlers()

See the module docs on ForgeOpsTracker.LoggerHandler for what install_handlers/0 actually covers (every process crash anywhere in the whole BEAM VM, via :logger: a meaningfully stronger automatic-capture story than any single-thread uncaught-exception hook) and capture_exception/3 for reporting one you've already rescued yourself.

Summary

Functions

Records one entry into the calling process's own breadcrumb trail: a query, an outbound call, or anything worth remembering right up to the moment something actually goes wrong. category defaults to "custom" and level to "info". A no-op, not a raise, when Configuration.t/0's own track_breadcrumbs is false.

Report an exception (or any raised/thrown term: see ForgeOpsTracker.EventBuilder's own doc for why this doesn't require a real Exception struct) you've already rescued yourself

Records one infrastructure reading (CPU, memory, disk, anything else a script of yours reads) from one of your own hosts. hostname defaults to Configuration.server_name, so a script running on the box it reports about needs no argument. Same buffered-batch delivery and no-op-when-disabled contract as capture_metric/2. A short-lived cron script that captures a few readings and lets the VM stop normally has them flushed on the supervised shutdown; call flush_metrics/0 if it might exit another way (System.halt/1).

Records a named business metric (a signup, a payment, anything you want to name), buffered and flushed periodically as one batch rather than one network call per capture. value defaults to 1.0 so a bare counter-style call needs no argument; pass one for a real magnitude (capture_metric("payment", 49.0)); it may be negative (a refund). A no-op when the client isn't enabled (no DSN, or this environment isn't in enabled_environments). Nothing here is automatic, so there is no track_metrics option.

Clears the calling process's own breadcrumb trail. Unlike Phoenix's own request lifecycle (which starts a genuinely fresh BEAM process per request, so nothing ever needs clearing between requests there), a long-lived process that handles more than one logical "unit of work" on the same pid, a GenServer processing several casts, a LiveView socket across several events, has to call this itself between units of work, or breadcrumbs from an earlier one will bleed into a later one's own report.

The calling process's current W3C trace id (32 lowercase hex characters), or nil outside a trace. Handy for your own logs: it's the id ForgeOps links errors across services with.

Ends the calling process's trace, sending it when the root took at least trace_capture_threshold milliseconds or an error was captured inside it. started_at is a DateTime, duration_ms a number.

Delivers every buffered metric and infrastructure reading right now, instead of waiting for the next flush interval.

Makes one outgoing HTTP call inside the calling process's trace: records it as an "http" span named "<METHOD> <host>" (never the path or query) and calls fun with the headers to add to the request, a list of {name, value} tuples holding a W3C traceparent whose parent id is that span's own id, so the called service's root span nests under it. Returns what fun returns; the span is recorded even if fun raises, which propagates unchanged.

Configures the client. Call once at startup: application.ex's own start/2, before the rest of your supervision tree starts, is the natural place. Accepts the same keys as ForgeOpsTracker.Configuration's struct fields

Installs the :logger handler described in ForgeOpsTracker.LoggerHandler's own module doc. Call once, after init/1. Safe to call more than once: a later call is a no-op.

Records one change to your system (a feature flag flipped, a config edit, a migration run) so ForgeOps can show it next to the errors and slowdowns that followed. Queued like an error report, so it never blocks the caller, and never raises; a no-op when the client isn't enabled.

Records one occurrence of transaction_name (a request, a database query, a queue job) taking duration_ms, bucketed by kind ("controller"/"job"/"query"). Not something app code normally calls directly: ForgeOpsTracker.Integrations.Phoenix/Ecto/Oban all call this once wired up via their own attach/0/attach/1.

Records a span you timed yourself under the current one; a no-op outside a trace. started_at is a DateTime.

record_span/5 with the same :statement and :db_system options as span/5.

Names the calling process's current request: transaction_name is the same name its performance sample and root span use, endpoint the HTTP method plus the route pattern ("GET /users/:id"), never the literal path. ForgeOpsTracker.Integrations.Phoenix does this as soon as the router has matched the route; a no-op outside a trace.

Manually attaches an affected user to whatever gets reported for the rest of this process (a Phoenix request's own connection process, a GenServer, an IEx session): stored in the process dictionary, the same "one process per request" isolation boundary Phoenix itself already gives every request, and the direct Elixir analog to Thread.current in gems/forge_ops_tracker's own equivalent. id/email/username are all independently optional; call with an empty map (or nothing set at all) to clear whatever was set.

Times fun as a child span of whatever span is open in this process (or of the request's root), returning its result. Outside a trace it just runs fun. Recorded even if fun raises, and the error propagates unchanged. kind is one of "controller", "service", "database", "redis", "http", "job", "other" (anything else is sent as "other").

span/4 with options. For a "database" span, :statement (the SQL) and :db_system (e.g. "postgresql") add "db.statement", with every string and number replaced by ?, and "db.system" to the span's data. The unmasked statement is never sent

Starts a trace on the calling process. ForgeOpsTracker.Integrations.Phoenix/Oban do this automatically; call it yourself (with finish_trace/4) to trace anything else, e.g. a GenServer handling a message. Pass a W3C traceparent header value (a job carrying the header it was enqueued with, say) to continue that trace: same trace id, and the root span's parent is the caller's span; a missing or malformed value starts a new trace.

Functions

add_breadcrumb(message, category \\ "custom", level \\ "info", data \\ %{})

@spec add_breadcrumb(String.t(), String.t(), String.t(), map()) :: :ok

Records one entry into the calling process's own breadcrumb trail: a query, an outbound call, or anything worth remembering right up to the moment something actually goes wrong. category defaults to "custom" and level to "info". A no-op, not a raise, when Configuration.t/0's own track_breadcrumbs is false.

Stored in the process dictionary, the same isolation boundary set_user/1 already uses and for the identical reason: a Phoenix request's own connection process, a GenServer, an Oban job all get their own separate trail, correctly isolated from every other process's.

capture_exception(reason, stacktrace, context \\ %{}, user \\ :from_process)

@spec capture_exception(term(), Exception.stacktrace(), map(), map() | nil) :: :ok

Report an exception (or any raised/thrown term: see ForgeOpsTracker.EventBuilder's own doc for why this doesn't require a real Exception struct) you've already rescued yourself:

try do
  charge_card(order)
rescue
  e ->
    ForgeOpsTracker.capture_exception(e, __STACKTRACE__, %{order_id: order.id})
    reraise e, __STACKTRACE__
end

stacktrace has to come from the actual rescue/catch site (__STACKTRACE__): unlike languages where an exception object carries its own trace, Elixir's stacktrace is only available via that special form, and only for as long as nothing else has run since the rescue/catch (the same reason the example above captures it before doing anything else).

user, if not given, defaults to whatever set_user/1 last set for this process (nil if nothing did); pass one explicitly to override that for this one report. The calling process's own breadcrumb trail (see add_breadcrumb/4) is always attached automatically, the same way user already is, with no separate argument for it.

Inside a trace (a Phoenix request, an Oban job, or your own start_trace/1), the event also carries the trace's trace_id, plus the request's transaction_name and endpoint under Phoenix, and the trace is marked errored so it's sent however fast it was.

capture_infrastructure_metric(name, value, hostname \\ nil)

@spec capture_infrastructure_metric(String.t(), number(), String.t() | nil) :: :ok

Records one infrastructure reading (CPU, memory, disk, anything else a script of yours reads) from one of your own hosts. hostname defaults to Configuration.server_name, so a script running on the box it reports about needs no argument. Same buffered-batch delivery and no-op-when-disabled contract as capture_metric/2. A short-lived cron script that captures a few readings and lets the VM stop normally has them flushed on the supervised shutdown; call flush_metrics/0 if it might exit another way (System.halt/1).

capture_metric(name, value \\ 1.0)

@spec capture_metric(String.t(), number()) :: :ok

Records a named business metric (a signup, a payment, anything you want to name), buffered and flushed periodically as one batch rather than one network call per capture. value defaults to 1.0 so a bare counter-style call needs no argument; pass one for a real magnitude (capture_metric("payment", 49.0)); it may be negative (a refund). A no-op when the client isn't enabled (no DSN, or this environment isn't in enabled_environments). Nothing here is automatic, so there is no track_metrics option.

clear_breadcrumbs()

@spec clear_breadcrumbs() :: :ok

Clears the calling process's own breadcrumb trail. Unlike Phoenix's own request lifecycle (which starts a genuinely fresh BEAM process per request, so nothing ever needs clearing between requests there), a long-lived process that handles more than one logical "unit of work" on the same pid, a GenServer processing several casts, a LiveView socket across several events, has to call this itself between units of work, or breadcrumbs from an earlier one will bleed into a later one's own report.

current_trace_id()

@spec current_trace_id() :: String.t() | nil

The calling process's current W3C trace id (32 lowercase hex characters), or nil outside a trace. Handy for your own logs: it's the id ForgeOps links errors across services with.

finish_trace(name, kind \\ "controller", started_at, duration_ms)

@spec finish_trace(String.t(), String.t(), DateTime.t(), number()) :: :ok

Ends the calling process's trace, sending it when the root took at least trace_capture_threshold milliseconds or an error was captured inside it. started_at is a DateTime, duration_ms a number.

flush_metrics()

@spec flush_metrics() :: :ok

Delivers every buffered metric and infrastructure reading right now, instead of waiting for the next flush interval.

http_span(method, url, data \\ %{}, fun)

@spec http_span(
  String.t() | atom(),
  String.t() | URI.t(),
  map(),
  (ForgeOpsTracker.Tracing.headers() ->
     result)
) :: result
when result: var

Makes one outgoing HTTP call inside the calling process's trace: records it as an "http" span named "<METHOD> <host>" (never the path or query) and calls fun with the headers to add to the request, a list of {name, value} tuples holding a W3C traceparent whose parent id is that span's own id, so the called service's root span nests under it. Returns what fun returns; the span is recorded even if fun raises, which propagates unchanged.

ForgeOpsTracker.http_span("POST", url, fn headers ->
  Req.post!(url, json: order, headers: headers)
end)

The headers are [] outside a trace, when propagate_traces is off, or when the host isn't in trace_propagation_targets; outside a trace no span is recorded either. With track_tracing off the header is still handed over (the trace id links errors across services) but no span is kept.

init(opts \\ [])

Configures the client. Call once at startup: application.ex's own start/2, before the rest of your supervision tree starts, is the natural place. Accepts the same keys as ForgeOpsTracker.Configuration's struct fields:

ForgeOpsTracker.init(dsn: "...", environment: "production")

Fields not passed here fall back to Application.get_env(:forge_ops_tracker, key) / FORGE_OPS_DSN-style environment variables read when this app started: see ForgeOpsTracker.Configuration's own module doc.

Once the client is enabled, this also sends a change snapshot (once per VM, from its own Task, so it never delays startup): see record_change/3 and the detect_changes option.

install_handlers()

@spec install_handlers() :: :ok

Installs the :logger handler described in ForgeOpsTracker.LoggerHandler's own module doc. Call once, after init/1. Safe to call more than once: a later call is a no-op.

record_change(kind, title, opts \\ [])

@spec record_change(atom() | String.t(), String.t(), keyword()) :: :ok

Records one change to your system (a feature flag flipped, a config edit, a migration run) so ForgeOps can show it next to the errors and slowdowns that followed. Queued like an error report, so it never blocks the caller, and never raises; a no-op when the client isn't enabled.

ForgeOpsTracker.record_change(:feature_flag, "Enabled new checkout",
  details: %{flag: "new_checkout", rollout: 25},
  actor: "alice@example.com"
)

kind is one of :feature_flag, :config, :migration, :dependency, :infrastructure, :other (atom or string; anything else is sent as "other"). title is cut to 200 characters. Options: :details (a small map), :environment (defaults to the configured one), :service, :actor, :url, :id (an idempotency key), :occurred_at (a DateTime or ISO 8601 string; defaults to now).

Separately, init/1 sends a change snapshot once per VM (runtime, loaded application versions, and, with track_env_var_names: true, environment variable names but never values), and ForgeOps records what changed since the last one. Turn that off with detect_changes: false.

record_performance(transaction_name, kind \\ "controller", duration_ms)

@spec record_performance(String.t(), String.t(), number()) :: :ok

Records one occurrence of transaction_name (a request, a database query, a queue job) taking duration_ms, bucketed by kind ("controller"/"job"/"query"). Not something app code normally calls directly: ForgeOpsTracker.Integrations.Phoenix/Ecto/Oban all call this once wired up via their own attach/0/attach/1.

Re-checks Configuration.enabled?/1 and track_performance on every call, the same "guarded every time, not just once at attach time" posture capture_exception/3 already has over Reporter.report/3: config can change at runtime, so a :telemetry.attached handler that checked this only once at startup could keep recording long after either flag turned off.

record_span(name, kind, started_at, duration_ms, data \\ %{})

@spec record_span(String.t(), String.t(), DateTime.t(), number(), map()) :: :ok

Records a span you timed yourself under the current one; a no-op outside a trace. started_at is a DateTime.

record_span(name, kind, started_at, duration_ms, data, opts)

@spec record_span(String.t(), String.t(), DateTime.t(), number(), map(), keyword()) ::
  :ok

record_span/5 with the same :statement and :db_system options as span/5.

set_request_route(transaction_name, endpoint)

@spec set_request_route(String.t() | nil, String.t() | nil) :: :ok

Names the calling process's current request: transaction_name is the same name its performance sample and root span use, endpoint the HTTP method plus the route pattern ("GET /users/:id"), never the literal path. ForgeOpsTracker.Integrations.Phoenix does this as soon as the router has matched the route; a no-op outside a trace.

set_user(user \\ %{})

@spec set_user(map()) :: :ok

Manually attaches an affected user to whatever gets reported for the rest of this process (a Phoenix request's own connection process, a GenServer, an IEx session): stored in the process dictionary, the same "one process per request" isolation boundary Phoenix itself already gives every request, and the direct Elixir analog to Thread.current in gems/forge_ops_tracker's own equivalent. id/email/username are all independently optional; call with an empty map (or nothing set at all) to clear whatever was set.

span(name, kind \\ "service", data \\ %{}, fun)

@spec span(String.t(), String.t(), map(), (-> result)) :: result when result: var

Times fun as a child span of whatever span is open in this process (or of the request's root), returning its result. Outside a trace it just runs fun. Recorded even if fun raises, and the error propagates unchanged. kind is one of "controller", "service", "database", "redis", "http", "job", "other" (anything else is sent as "other").

order = ForgeOpsTracker.span("charge card", "service", %{order_id: id}, fn -> charge(id) end)

A "database" span can carry its SQL with span/5; see there.

span(name, kind, data, opts, fun)

@spec span(String.t(), String.t(), map(), keyword(), (-> result)) :: result
when result: var

span/4 with options. For a "database" span, :statement (the SQL) and :db_system (e.g. "postgresql") add "db.statement", with every string and number replaced by ?, and "db.system" to the span's data. The unmasked statement is never sent:

sql = "SELECT id, total FROM orders WHERE customer_id = $1 AND status = 'paid'"

rows =
  ForgeOpsTracker.span("Load orders", "database", %{}, [statement: sql, db_system: "postgresql"], fn ->
    Repo.query!(sql, [customer_id]).rows
  end)

Other options, and these two on any other kind, are ignored.

start_trace(traceparent \\ nil)

@spec start_trace(String.t() | nil) :: :ok

Starts a trace on the calling process. ForgeOpsTracker.Integrations.Phoenix/Oban do this automatically; call it yourself (with finish_trace/4) to trace anything else, e.g. a GenServer handling a message. Pass a W3C traceparent header value (a job carrying the header it was enqueued with, say) to continue that trace: same trace id, and the root span's parent is the caller's span; a missing or malformed value starts a new trace.

A trace id exists whenever the client is enabled, even with track_tracing off, since errors captured inside the trace carry it and http_span/4 propagates it; only span recording is gated on track_tracing.