LangExtract (LangExtract v0.6.0)

Copy Markdown View Source

Extracts structured data from text with source grounding. Maps extraction strings back to exact byte positions in source text.

This module is the main entry point: new/2 builds a client, run/4 executes the full pipeline, and align/3 / extract/3 expose the lower-level steps. Beyond the facade:

Summary

Functions

Aligns extraction strings to byte spans in source text.

Parses LLM output, aligns extractions against source text, and returns enriched spans with class and attributes.

Creates a configured LLM client for extraction.

Runs the full extraction pipeline: prompt → LLM → parse → align.

Types

provider()

@type provider() :: :claude | :openai | :gemini

Functions

align(source, extractions, opts \\ [])

@spec align(String.t(), [String.t()], keyword()) :: [LangExtract.Alignment.Span.t()]

Aligns extraction strings to byte spans in source text.

Returns a list of %LangExtract.Alignment.Span{} structs, one per extraction.

Options

  • :fuzzy_threshold - minimum overlap ratio for fuzzy match (default 0.75)

Examples

iex> LangExtract.align("the quick brown fox", ["quick brown"])
[%LangExtract.Alignment.Span{text: "quick brown", byte_start: 4, byte_end: 15, status: :exact}]

extract(source, raw_llm_output, opts \\ [])

@spec extract(String.t(), String.t(), keyword()) ::
  {:ok, [LangExtract.Alignment.Span.t()]}
  | {:error, {:invalid_format, String.t()} | :missing_extractions}

Parses LLM output, aligns extractions against source text, and returns enriched spans with class and attributes.

Accepts both canonical and dynamic-key format (where each entry uses the class name as the key), in JSON or YAML. Strips markdown fences and think tags before parsing.

Options

  • :fuzzy_threshold - minimum overlap ratio for fuzzy match (default 0.75)

Examples

iex> raw = ~s({"extractions": [{"class": "word", "text": "fox"}]})
iex> {:ok, [span]} = LangExtract.extract("the quick brown fox", raw)
iex> span.status
:exact

new(provider, opts \\ [])

@spec new(
  provider(),
  keyword()
) :: LangExtract.Client.t()

Creates a configured LLM client for extraction.

Raises ArgumentError on an unknown provider or unbuildable HTTP client (e.g. missing API key). The raise is deliberate: misconfiguration here is a programmer error caught at client construction, while runtime failures during extraction (run/4, extract/3) return tagged tuples.

Examples

client = LangExtract.new(:claude, api_key: "sk-...")
client = LangExtract.new(:openai, api_key: "sk-...", model: "gpt-4o")
client = LangExtract.new(:gemini, api_key: "gm-...")

run(client, source, template, opts \\ [])

Runs the full extraction pipeline: prompt → LLM → parse → align.

Returns {:ok, {spans, chunk_errors}} on success. When some chunks fail to parse, the successful spans are still returned alongside the errors. Returns {:error, reason} only for infrastructure failures (task exits, timeouts).

Options

  • :max_chunk_chars - chunk size in characters (default 1000)
  • :max_concurrency - parallel chunk requests (default 10)
  • :task_timeout - per-chunk task timeout (default :infinity)
  • :fuzzy_threshold - minimum LCS coverage for fuzzy match (default 0.75)
  • :min_density - minimum matched-token density of a fuzzy span (default 1/3)
  • :accept_lesser - allow prefix-fragment grounding (default true)
  • :exact_algorithm - :dp (occurrence DP, default) or :first_occurrence

See the "Alignment and Spans" guide for what the alignment options tune.

Examples

client = LangExtract.new(:claude, api_key: "sk-...")
template = %LangExtract.Prompt.Template{description: "Extract entities."}
{:ok, {spans, errors}} = LangExtract.run(client, "the quick brown fox", template)