Imp.Predict.KNN (Imp v0.5.0)

Copy Markdown View Source

An embedding-based nearest-neighbor retriever over a trainset, with the same query and ranking semantics as DSPy 3.2.1 KNN.

Construction embeds every trainset example ONCE through the required :vectorizer (any Imp.Embeddings provider — a module implementing embed/2 or a 2-arity function; DSPy's Embedder analog). Each example is rendered as "key: value" pairs for its INPUT fields joined by " | ", exactly like upstream's trainset_casted_to_vectorize. At call time the query inputs are rendered the same way, embedded, and scored against the trainset by dot product; the top k examples return in descending-score order (upstream's argsort()[-k:][::-1]).

Two details matter when moving data between Python and Elixir:

  • Field iteration order: Python dicts iterate in insertion order; Elixir maps do not preserve insertion order. Pass keyword-list inputs (or single-input-field examples) when the exact multi-field rendering order matters for embedding equality.
  • Tie-breaking among equal scores follows a stable sort (higher original index wins within a tie, matching a stable argsort); NumPy's default quicksort leaves ties unspecified.

For token-overlap retrieval without an embedding provider, use Imp.Retrieve.Memory.

Summary

Functions

Returns the k nearest trainset examples for the given inputs (DSPy KNN.__call__): embed the query, dot-product against the trainset vectors, take the top k by descending score.

Builds the retriever, embedding the trainset once (DSPy KNN.__init__).

Functions

call(knn, inputs)

Returns the k nearest trainset examples for the given inputs (DSPy KNN.__call__): embed the query, dot-product against the trainset vectors, take the top k by descending score.

new(k, trainset, opts \\ [])

Builds the retriever, embedding the trainset once (DSPy KNN.__init__).

Options:

  • :vectorizer (required) — an Imp.Embeddings provider: a module exporting embed/2 or a function of (texts, opts).