What one scan DID — the poll and the drain it triggered, as a single value.
A scan is two phases that always happen together and were reported
separately. The poll says what it fetched, what it could not reach and what it
cost; the drain returns a ReactiveDag.Drain.Report saying what recomputed
downstream. A host wanting "what happened when this ran" had to collect both
from one telemetry payload and know which half answered which question.
This is the pair, named. ReactiveDag.ScanWorker builds one per scan and puts
it on [:reactive_dag, :scan, :stop], so a broadcast, a durable row or a log
line takes one value instead of reassembling it.
ReactiveDag.Insights.record/1 retains it whole for the same reason — a host
unwrapping to run.report throws away the poll, which is most of what the
scan did and all of what it could not reach.
It does not merge the phases
Deliberately. A poll is fallible network I/O that must not abort a sweep when one upstream is down; a drain is pure recomputation that must RAISE rather than mark keys clean over work that did not happen. Those two rules cannot live in one loop, and nothing here tries to make them.
What was genuinely duplicated was the reporting: two vocabularies for "what
did this cost", answered by ReactiveDag.Rollup for both. This is the same
move one level up — one value for "what did this run do".
The drain may be absent
report is nil for a scan that never drained: an unscannable source (no
credential, integration not enabled) is a completed scan that found nothing,
and a host recording scan results still wants the row. drained?/1 says
which, rather than making every caller test for nil.
Summary
Types
The poll half — what ReactiveDag.Source.refresh/3 returned, plus the cell it
belongs to.
Functions
One cost key across both phases, summed per bucket.
Did the poll find anything?
Could the poll see everything it meant to?
Did a drain run at all?
Sum one cost key across BOTH phases.
Types
@type t() :: %ReactiveDag.ScanRun{ cell: String.t() | :sweep, changed: [String.t()], detail: map(), duration_us: non_neg_integer(), not_scannable: term() | nil, report: ReactiveDag.Drain.Report.t() | nil, unreachable: [{String.t(), term()}] }
The poll half — what ReactiveDag.Source.refresh/3 returned, plus the cell it
belongs to.
Functions
One cost key across both phases, summed per bucket.
ScanRun.by(run, :tokens_in)
#=> %{"claude-haiku-4-5" => 900, "openai/gpt-5.6-luna" => 3000}The breakdown behind total/2, and the only form that answers "which model is
driving this" when the poll and the drain use different ones — which is the
normal case, since a classifier and a summariser are chosen separately.
Did the poll find anything?
Distinct from drained?/1: a poll can change keys whose recompute produced
nothing downstream, and a drain can run over marks another source left.
Could the poll see everything it meant to?
The honest-gap discipline in one predicate: a scan that could not look must never render as a scan that found nothing.
Did a drain run at all?
false for an unscannable source, which completes without draining. A host
rendering "0 passes" for that would be reporting a drain that never happened.
Sum one cost key across BOTH phases.
ScanRun.total(run, :tokens_in)This is the question a scan's cost line actually asks, and until now it took
two calls and an addition: the crawl's own spend lives in detail and its
downstream recomputes' spend lives in the report's steps. A crawler that
classifies documents with a model and feeds nodes that summarise them with
another spends in both places, and neither number alone is the bill.
ReactiveDag.Rollup does the arithmetic, so a flat count and a per-model
breakdown both total, mixed freely across the two phases.