The find surface — two composable layers on one call.
- Filters — tag-AND and derived source, fetched through
YmerNode.References.list_classified/1rather than from the repo directly, so those rules stay implemented once. The resulting cycle between the two modules is runtime-only: adefdelegateout and a plain call back. - Keyword floor — a score of distinct query terms matched against
YmerNode.References.Reference.search_text/1, which is every text field, so a term carried only by a uri still matches.
mode names what ran: :keyword when a query was given, :filter when none
was — filters only, newest first. There are two modes and not three. Ranking
by embedding is a knowing omission rather than an unimplemented branch: the
node ships no embedding provider for the registry, and a mode that silently
fell back would tell a caller its query had been ranked one way when it had
been ranked the other.
Ties break newest-updated first, then by id descending, so an equal-scoring pair has a stable order rather than the repo's.
Summary
Functions
Runs a find. Options: :query (a string, or nil for filter mode), :tags
(tag-AND list), :source (a source token), :limit (a positive integer, or
nil for all).
Splits a query into the distinct terms a keyword score counts.
Functions
Runs a find. Options: :query (a string, or nil for filter mode), :tags
(tag-AND list), :source (a source token), :limit (a positive integer, or
nil for all).
Answers %{mode: :keyword | :filter, results: results}, where each result is
%{reference: %YmerNode.References.Reference{}, classification: %{source: source, recipe: recipe}, score: number | nil} — the score is nil in filter
mode, because nothing ranked.
The classification rides along rather than being derived by whoever renders
the hit: YmerNode.References.list_classified/1 had to compute it to apply
the source filter, and passing :declarations through means one operation
reads the seam once. :declarations is optional here for the same reason it
is there — a caller with none in hand gets one read rather than a crash.
Splits a query into the distinct terms a keyword score counts.
Downcased, split on anything that is not a letter, a number or a hyphen — so
an identifier like ABC-1234 survives as one term — then deduplicated and
capped.
Examples
iex> YmerNode.References.Search.tokenize("Release notes, ABC-1234")
["release", "notes", "abc-1234"]
iex> YmerNode.References.Search.tokenize("a I x")
[]
iex> YmerNode.References.Search.tokenize("dup DUP dup")
["dup"]