The diffing logic behind mix corpus.coverage (px-35i.4 Phase 6).
The corpus is authored JSON, not extracted from the existing 16k-line ExUnit suite (see the plan's "The generation mechanism" section) - so nothing keeps the corpus in step with what that suite actually exercises. This module is the coverage instrument that closes that loop: it turns the suite into an authoring checklist rather than a source, without making it the corpus's source of truth.
What this measures, and what it does not
This is a heuristic, best-effort report, not a precise instrument. It
works in three steps, all pure functions here plus real compiler calls -
never eval, never hand-rolled opcode matching against source text:
extract_sources/1statically scans one.exsfile's AST for calls toPredicator.evaluate/2,Predicator.evaluate/3, orPredicator.compile/1whose first argument is a literal string, and returns those literal source strings. This is exactly the ~25% static-AST pass the plan's "Current State Analysis" section already describes and sanctions as a legitimate, partial signal - used here for a report, not as the corpus's authoring mechanism, so the drift caveats that ruled it out as the mechanism do not apply.suite_pattern_frequencies/1compiles each extracted source through the realPredicator.compile/1(not string matching) and counts patterns: an opcode name for most instructions, but"compare:OPERATOR"forcompareand"call:function_name"forcall, because those two opcodes carry the operand that actually distinguishes one case from another (STRICT_EQvsGT,lenvsupper) - the same operand-aware cutPredicator.Conformance.Featuresalready makes for feature tags. A source that fails to compile is skipped; the AST scan is approximate and a false-positive extraction (e.g. a string that merely looks like a literal source but is test scaffolding) is expected and harmless here.corpus_pattern_frequencies/1computes the same patterns from the shipped corpus's owninstructions(already decoded JSON case maps), anddiff/2reports every pattern the suite's static extraction hits more often than the corpus does, grouped by tier.
What this does not see: the ~75% of suite assertions that build context
from a bound variable or executed setup (unresolvable without running the
suite), any assertion carrying a third opts argument, error-shape
assertions, or instruction-list literals fed directly to evaluate/2,3.
Those are exactly the classes the bead's own extractability note names as
unreachable by a static pass. The report is therefore a floor on suite
coverage, not a ceiling, and its top entries need a human read before being
authored as cases - some will be Elixir-internal scaffolding (opts
handling, struct internals) rather than genuine corpus gaps.
Classifying tier-5 (call:<name>) gaps
diff/2 alone cannot tell a genuine builtin gap from noise: a call:<name>
pattern's name might be a test-registered function (boom, double, ...)
that a sibling implementation could never implement, or a builtin whose
result is non-deterministic (Date.now, Math.random) and therefore can
never get a pinned corpus case either. classify/2 sorts every gap into
one of three statuses (px-q1f):
:gap- a genuine builtin gap, worth authoring.:excluded- a documented exclusion:relative_date(an opcode, seeconformance/README.md's "Opcodes excluded from the coverage rule") or a name indocumented_exclusion_functions/0(see the README's "Functions excluded from the coverage rule"). Shown inline with a note rather than as a barecorpus: 0row.:suite_local- acall:<name>whose name is not a key of the caller- supplied builtin registry (in practicePredicator.Evaluator. merge_functions([])), so it was registered by an individual ExUnit test via its ownopts[:functions]map and can never become a corpus case.format_report/1moves these to a trailing labelled section rather than dropping them silently - the report is a heuristic instrument, and a reader who remembers writing a test for one of these names should be able to see where it went rather than wonder if the scan missed it.
Classification is deliberately not a hardcoded name list on the "this is
a builtin" side: classify/2 takes the registry as data from the caller
(mix corpus.coverage passes Predicator.Evaluator.merge_functions([])'s
keys) so a new builtin needs no update here to stop showing up as
suite-local.
Report only: nothing here writes a case or fails a gate.
Summary
Types
One coverage gap: a pattern the suite hits more than the corpus does.
A gap's classification, set by classify/2 (absent - equivalently :gap -
on a gap fresh out of diff/2, which knows nothing about the builtin
registry)
A pattern: an opcode name, or "opcode:operand" for compare/call.
pattern => how many times it was observed.
Functions
Classifies each gap's :status (see the moduledoc's "Classifying tier-5
gaps" section and gap_status/0): :gap, :excluded, or :suite_local.
Computes pattern frequencies from decoded corpus case maps (the shipped
conformance/corpus/tier-*.json cases, already JSON.decode/1-ed).
Diffs suite pattern frequencies against corpus pattern frequencies.
The builtin function names documented as non-deterministic exclusions in
conformance/README.md's "Functions excluded from the coverage rule"
section - Date.now and Math.random. classify/2 marks a call:<name>
gap whose name is in this set as :excluded rather than :gap, the same
treatment relative_date gets as an opcode-level exclusion.
Extracts literal source strings from one file's Elixir source.
Formats a list of gaps (as returned by diff/2, typically classified by
classify/2 first) as a human-readable report, grouped by tier so it is
actionable one tier at a time.
Computes pattern frequencies for a list of extracted source strings.
Types
@type gap() :: %{ :tier => pos_integer(), :tier_name => String.t(), :pattern => pattern(), :suite_count => pos_integer(), :corpus_count => non_neg_integer(), optional(:status) => gap_status() }
One coverage gap: a pattern the suite hits more than the corpus does.
@type gap_status() :: :gap | :excluded | :suite_local
A gap's classification, set by classify/2 (absent - equivalently :gap -
on a gap fresh out of diff/2, which knows nothing about the builtin
registry):
:gap- worth authoring.:excluded- a documented, non-deterministic exclusion; see the moduledoc's "Classifying tier-5 gaps" section.:suite_local- a test-registered function name, never a corpus candidate.
@type pattern() :: String.t()
A pattern: an opcode name, or "opcode:operand" for compare/call.
@type pattern_frequencies() :: %{required(pattern()) => pos_integer()}
pattern => how many times it was observed.
Functions
Classifies each gap's :status (see the moduledoc's "Classifying tier-5
gaps" section and gap_status/0): :gap, :excluded, or :suite_local.
builtin_names is the set of function names the classifier treats as
real builtins - in practice the caller passes
Predicator.Evaluator.merge_functions([])'s keys, so this module never
hardcodes "what is a builtin" itself. A gap whose pattern is not a
call:<name> pattern (i.e. not tier 5) is :gap unless it is the
documented opcode exclusion relative_date.
Examples
iex> gaps = [
...> %{tier: 5, tier_name: "functions", pattern: "call:Math.abs", suite_count: 3, corpus_count: 0},
...> %{tier: 5, tier_name: "functions", pattern: "call:Date.now", suite_count: 6, corpus_count: 0},
...> %{tier: 5, tier_name: "functions", pattern: "call:boom", suite_count: 2, corpus_count: 0},
...> %{tier: 4, tier_name: "rich types", pattern: "relative_date", suite_count: 4, corpus_count: 0}
...> ]
iex> Predicator.Conformance.Coverage.classify(gaps, MapSet.new(["Math.abs", "Date.now"]))
...> |> Enum.map(& &1.status)
[:gap, :excluded, :suite_local, :excluded]
@spec corpus_pattern_frequencies([map()]) :: pattern_frequencies()
Computes pattern frequencies from decoded corpus case maps (the shipped
conformance/corpus/tier-*.json cases, already JSON.decode/1-ed).
Reads each case's "instructions" field the same way suite_pattern_frequencies/1
reads a compiled instruction list - the tagged-value codec only affects
operand values, never the opcode strings this counts.
Examples
iex> cases = [%{"instructions" => [["load", "a"], ["lit", 1], ["compare", "GT"]]}]
iex> Predicator.Conformance.Coverage.corpus_pattern_frequencies(cases)
%{"load" => 1, "lit" => 1, "compare:GT" => 1}
@spec diff(pattern_frequencies(), pattern_frequencies()) :: [gap()]
Diffs suite pattern frequencies against corpus pattern frequencies.
Returns every pattern the suite hits more than the corpus does (corpus
absence counts as 0), grouped implicitly by carrying each entry's tier -
format_report/1 does the grouping for display. A pattern whose opcode is
not in Predicator.Instructions.opcodes/0 (should not happen for anything
that compiled, but kept defensive) is skipped rather than crashing the
report.
Sorted by tier, then by descending suite count, then by pattern name, so the most-exercised gap in the lowest tier sorts first - the tier a sibling is most likely to be working on right now.
Examples
iex> suite = %{"compare:GT" => 5, "compare:STRICT_EQ" => 2, "lit" => 10}
iex> corpus = %{"compare:GT" => 5, "lit" => 3}
iex> Predicator.Conformance.Coverage.diff(suite, corpus)
[
%{tier: 1, tier_name: "core", pattern: "lit", suite_count: 10, corpus_count: 3},
%{tier: 1, tier_name: "core", pattern: "compare:STRICT_EQ", suite_count: 2, corpus_count: 0}
]
The builtin function names documented as non-deterministic exclusions in
conformance/README.md's "Functions excluded from the coverage rule"
section - Date.now and Math.random. classify/2 marks a call:<name>
gap whose name is in this set as :excluded rather than :gap, the same
treatment relative_date gets as an opcode-level exclusion.
Examples
iex> Predicator.Conformance.Coverage.documented_exclusion_functions() |> Enum.sort()
["Date.now", "Math.random"]
Extracts literal source strings from one file's Elixir source.
Scans the AST (not a regex over the text) for calls to
Predicator.evaluate/2, Predicator.evaluate/3, or Predicator.compile/1
whose first argument is a literal string. A file that fails to parse
returns [] rather than raising - this is a best-effort scan over test
files, not a compiler.
Examples
iex> Predicator.Conformance.Coverage.extract_sources("Predicator.evaluate(\"a > 1\", %{})")
["a > 1"]
iex> Predicator.Conformance.Coverage.extract_sources("Predicator.compile(\"x\")")
["x"]
iex> Predicator.Conformance.Coverage.extract_sources("Predicator.evaluate(source, %{})")
[]
iex> Predicator.Conformance.Coverage.extract_sources("not valid elixir (")
[]
Formats a list of gaps (as returned by diff/2, typically classified by
classify/2 first) as a human-readable report, grouped by tier so it is
actionable one tier at a time.
A gap's :status (absent counts as :gap) controls how it is shown:
:gap prints plainly, :excluded gets an inline note instead of
appearing as an unqualified corpus: 0 row, and :suite_local entries are
pulled out of every tier and moved to a trailing labelled section rather
than being silently dropped - this is a heuristic instrument, and a reader
who remembers writing a test against one of these names should see where it
went, not wonder if the scan missed it.
Examples
iex> Predicator.Conformance.Coverage.format_report([])
"No coverage gaps found: every pattern the suite's extracted sources exercise\nappears at least as often in the corpus.\n"
@spec suite_pattern_frequencies([String.t()]) :: pattern_frequencies()
Computes pattern frequencies for a list of extracted source strings.
Each source is compiled through the real Predicator.compile/1. A source
that fails to compile contributes nothing - the extraction in
extract_sources/1 is approximate, so an uncompilable "source" is expected
(e.g. a string that merely looked like one).
Examples
iex> Predicator.Conformance.Coverage.suite_pattern_frequencies(["1 > 2", "1 > 2", "a == b"])
%{"compare:GT" => 2, "compare:EQ" => 1, "lit" => 4, "load" => 2}