An embedding-based nearest-neighbor retriever over a trainset, with the same
query and ranking semantics as DSPy 3.2.1 KNN.
Construction embeds every trainset example ONCE through the required
:vectorizer (any Imp.Embeddings provider — a module implementing
embed/2 or a 2-arity function; DSPy's Embedder analog). Each example is
rendered as "key: value" pairs for its INPUT fields joined by " | ",
exactly like upstream's trainset_casted_to_vectorize. At call time the
query inputs are rendered the same way, embedded, and scored against the
trainset by dot product; the top k examples return in descending-score
order (upstream's argsort()[-k:][::-1]).
Two details matter when moving data between Python and Elixir:
- Field iteration order: Python dicts iterate in insertion order; Elixir maps do not preserve insertion order. Pass keyword-list inputs (or single-input-field examples) when the exact multi-field rendering order matters for embedding equality.
- Tie-breaking among equal scores follows a stable sort (higher original
index wins within a tie, matching a stable
argsort); NumPy's default quicksort leaves ties unspecified.
For token-overlap retrieval without an embedding provider, use
Imp.Retrieve.Memory.
Summary
Functions
Returns the k nearest trainset examples for the given inputs
(DSPy KNN.__call__): embed the query, dot-product against the trainset
vectors, take the top k by descending score.
Builds the retriever, embedding the trainset once (DSPy KNN.__init__).
Functions
Returns the k nearest trainset examples for the given inputs
(DSPy KNN.__call__): embed the query, dot-product against the trainset
vectors, take the top k by descending score.
Builds the retriever, embedding the trainset once (DSPy KNN.__init__).
Options:
:vectorizer(required) — anImp.Embeddingsprovider: a module exportingembed/2or a function of(texts, opts).