Threadline.Health (Threadline v0.12.0)

Copy Markdown View Source

Health checks for Threadline infrastructure.

Queries the PostgreSQL system catalog to verify trigger installation status for all user tables.

Mix-task parity

See mix threadline.health.coverage for a viewer with --json, --schema=NAME, and --strict flags. Viewer by default (exits 0); --strict turns :error-severity findings into exit 1. The task does not exit non-zero on uncovered tables even with --strict; use mix threadline.verify_coverage for the positive-list CI gate.

Telemetry

On every successful call, trigger_coverage/1 emits [:threadline, :health, :checked] with measurements %{covered: integer, uncovered: integer, expected_uncovered: integer}. The expected_uncovered measurement is an additive — old subscribers reading only covered/uncovered keep working unchanged).

trigger_findings/1 emits [:threadline, :health, :findings_checked] with measurements %{errors: integer, warnings: integer}, counted over the findings list it returns. legacy_key_findings/1 emits no telemetry event.

Summary

Functions

Returns a list of Threadline.Health.Finding structs (:unresolved_legacy_keys) for audit rows captured before their table's trigger was regenerated and still carrying an unresolved primary key — rows history/3 cannot find by key. Unlike trigger_findings/1, which is catalog-only, this scans audit_changes per table.

Returns a list of tagged tuples indicating trigger coverage for all user tables in the given schema (default "public").

Returns a list of Threadline.Health.Finding structs describing detected capture problems: disabled or replica-only triggers, duplicate capture triggers, drifted or missing key columns, and shared per-table functions.

Functions

legacy_key_findings(opts)

@spec legacy_key_findings(keyword()) :: [Threadline.Health.Finding.t()]

Returns a list of Threadline.Health.Finding structs (:unresolved_legacy_keys) for audit rows captured before their table's trigger was regenerated and still carrying an unresolved primary key — rows history/3 cannot find by key. Unlike trigger_findings/1, which is catalog-only, this scans audit_changes per table.

Options

  • :repo — required Ecto.Repo module.
  • :schema — same as trigger_findings/1: a schema name string, or a list of schema name strings. Omitting it covers every non-system schema.
  • :statement_timeout — milliseconds, default 15_000. Applied with a transaction-local setting, so it is safe through PgBouncer transaction pooling. When the timeout elapses — typically a missing row-history index — this function raises Postgrex.Error with postgres code :query_canceled; see Step 4.

Each table's probe is capped at 10,000 rows; a capped finding's details["unresolved_count"] is 10000 and its message reads "at least 10000". DELETE rows, rows whose key columns were redacted or are otherwise absent from data_after, and dropped tables are never counted — see What cannot be recovered. A finding's details map has string keys "unresolved_count" (integer), "capped" (boolean), and "key_columns" (list of strings).

Emits no telemetry event.

Example

Threadline.Health.legacy_key_findings(repo: MyApp.Repo)
#=> [%Threadline.Health.Finding{code: :unresolved_legacy_keys, ...}]

trigger_coverage(opts)

Returns a list of tagged tuples indicating trigger coverage for all user tables in the given schema (default "public").

Audit tables (audit_transactions, audit_changes, audit_actions) are excluded from the result — they are not expected to have triggers (CAP-10).

A third tuple variant {:expected_uncovered, name} is supported for bookkeeping tables that are intentionally not audited (e.g. schema_migrations). The bucket is computed from a hardcoded baseline plus config :threadline, :health, expected_uncovered_tables: [...], with :audit_anyway removing entries from the union.

Options

  • :repo — required Ecto.Repo module
  • :schema — optional schema name string (default "public"). Programmatic callers are responsible for sanitizing or trusting their own input — this function does NOT validate :schema against pg_namespace. Surfaces that take untrusted input (LV / Mix task) MUST validate at the edge.

A disabled or replica-only trigger no longer counts as covered — see trigger_findings/1, which reports it as :capture_trigger_disabled.

Returns [{:covered | :uncovered | :expected_uncovered, table_name}].

Example

Threadline.Health.trigger_coverage(repo: MyApp.Repo)
#=> [{:covered, "users"}, {:expected_uncovered, "schema_migrations"}, {:uncovered, "orders"}]

trigger_findings(opts)

@spec trigger_findings(keyword()) :: [Threadline.Health.Finding.t()]

Returns a list of Threadline.Health.Finding structs describing detected capture problems: disabled or replica-only triggers, duplicate capture triggers, drifted or missing key columns, and shared per-table functions.

Options

  • :repo — required Ecto.Repo module.
  • :schema — a schema name string, or a list of schema name strings. Omitting it covers every non-system schema (excludes pg_catalog, information_schema, pg_toast*, pg_temp*, and the configured Threadline storage schema's own tables). This is deliberately different from trigger_coverage/1's "public" default: two tables with the same name in different schemas, and a per-table function shared across schemas, cannot be seen one schema at a time. The shared-function check always scans the whole catalog regardless of :schema — the option only filters which findings, by their table's schema, are returned.

A malformed :trigger_capture config raises the same ArgumentError that the internal trigger-capture config loader raises for capture itself.

Findings are sorted by {schema, table, code}, with message as the final tie-break, so two consecutive calls return identical lists. One table may produce more than one finding; there is no short-circuit.

Example

Threadline.Health.trigger_findings(repo: MyApp.Repo)
#=> [%Threadline.Health.Finding{code: :capture_trigger_disabled, ...}]