Extracts structured data from text with source grounding. Maps extraction strings back to exact byte positions in source text.
This module is the main entry point: new/2 builds a client,
template!/2 builds a validated task definition (template/2 is its
non-raising twin for runtime task data), run/4 executes the full
pipeline, and align/3 / extract/3 expose the lower-level steps.
Beyond the facade:
LangExtract.Prompt.Validator— pre-flight check that few-shot examples align against their own source textLangExtract.Serializer— convert results to plain maps and JSONL for storage or interopLangExtract.Extraction— the extraction struct used in template examples and parsed LLM output
Summary
Functions
Aligns extraction strings to byte spans in source text.
Parses LLM output, aligns extractions against source text, and returns enriched spans with class and attributes.
Creates a configured LLM client for extraction.
Runs the full extraction pipeline: prompt → LLM → parse → align.
Streams per-chunk extraction results as each chunk completes.
Builds a validated extraction template, returning a tagged tuple.
Builds a validated extraction template.
Types
Functions
@spec align(String.t(), [String.t()], keyword()) :: [LangExtract.Span.t()]
Aligns extraction strings to byte spans in source text.
Returns a list of %LangExtract.Span{} structs, one per extraction.
Options
:fuzzy_threshold- minimum overlap ratio for fuzzy match (default0.75)
Examples
iex> LangExtract.align("the quick brown fox", ["quick brown"])
[%LangExtract.Span{text: "quick brown", byte_start: 4, byte_end: 15, status: :exact}]
@spec extract(String.t(), String.t(), keyword()) :: {:ok, [LangExtract.Span.t()]} | {:error, {:invalid_format, String.t()} | :missing_extractions}
Parses LLM output, aligns extractions against source text, and returns enriched spans with class and attributes.
Accepts both canonical and dynamic-key format (where each entry uses the class name as the key). JSON only (since 0.7.0). Strips markdown fences and think tags before parsing.
Options
:fuzzy_threshold- minimum overlap ratio for fuzzy match (default0.75)
Examples
iex> raw = ~s({"extractions": [{"class": "word", "text": "fox"}]})
iex> {:ok, [span]} = LangExtract.extract("the quick brown fox", raw)
iex> span.status
:exact
@spec new( provider(), keyword() ) :: LangExtract.Client.t()
Creates a configured LLM client for extraction.
Raises ArgumentError on an unknown provider or unbuildable HTTP client
(e.g. missing API key). The raise is deliberate: misconfiguration here is
a programmer error caught at client construction, while runtime failures
during extraction (run/4, extract/3) return tagged tuples.
Examples
client = LangExtract.new(:claude, api_key: "sk-...")
client = LangExtract.new(:openai, api_key: "sk-...", model: "gpt-4o")
client = LangExtract.new(:gemini, api_key: "gm-...")
@spec run(LangExtract.Client.t(), String.t(), LangExtract.Template.t(), keyword()) :: {:ok, LangExtract.Result.t()} | {:error, {:task_exit, term()}}
Runs the full extraction pipeline: prompt → LLM → parse → align.
Returns {:ok, %Result{}} on success — document-ordered spans plus
per-chunk errors. When some chunks fail to parse, the successful spans
are still returned alongside the errors. Returns {:error, reason} only
for infrastructure failures (task exits, timeouts).
Not a drop-in swap with LangExtract.Runner.run/4 despite the matching
shape: the runner retries failures into per-chunk errors and never
returns {:error, _}, while this function abandons the document on a
task exit. See the failure-semantics table in the "Running in
Production" guide.
Options
:max_chunk_chars- chunk size in characters (default1000):max_concurrency- parallel chunk requests (default10):task_timeout- per-chunk task timeout (default:infinity):fuzzy_threshold- minimum LCS coverage for fuzzy match (default0.75):min_density- minimum matched-token density of a fuzzy span (default1/3):accept_lesser- allow prefix-fragment grounding (defaulttrue):exact_algorithm-:dp(occurrence DP, default) or:first_occurrence
Chunk size is measured in characters; span offsets are always bytes. See the "Alignment and Spans" guide for what the alignment options tune.
Examples
client = LangExtract.new(:claude, api_key: "sk-...")
template = LangExtract.template!("Extract entities.")
{:ok, %LangExtract.Result{spans: spans, errors: errors}} =
LangExtract.run(client, "the quick brown fox", template)
@spec stream(LangExtract.Client.t(), String.t(), LangExtract.Template.t(), keyword()) :: Enumerable.t()
Streams per-chunk extraction results as each chunk completes.
Returns a lazy stream of {:ok, %ChunkResult{}} and
{:error, %ChunkError{}} events in completion order, not
document order — consumers who need latency don't wait for slow chunks;
consumers who need order sort by the byte ranges every event carries.
Nothing runs until the stream is consumed, and a slow consumer naturally
limits how many chunk requests are in flight.
Failure semantics differ from run/4 deliberately: every failure stays
per-chunk. A chunk whose task times out is reported as
{:error, %ChunkError{reason: {:task_exit, :timeout}}} with its byte
range, and the remaining chunks keep flowing — where run/4 abandons
the document and returns {:error, {:task_exit, reason}}.
Takes the same options as run/4.
Examples
client
|> LangExtract.stream(document, template)
|> Enum.each(fn
{:ok, chunk_result} -> handle_spans(chunk_result.spans)
{:error, chunk_error} -> log_failure(chunk_error)
end)
@spec template( String.t(), keyword() ) :: {:ok, LangExtract.Template.t()} | {:error, Exception.t()}
Builds a validated extraction template, returning a tagged tuple.
The non-raising twin of template!/2 for templates built from runtime
data (user-uploaded or JSON-loaded task definitions). Returns
{:error, exception} where template!/2 would raise — an
ArgumentError for malformed example maps, or a
LangExtract.Prompt.Validator.ValidationError (carrying the per-example
issues) for examples whose extractions don't align.
Examples
iex> {:error, %ArgumentError{}} =
...> LangExtract.template("Extract.", examples: [%{extractions: []}])
@spec template!( String.t(), keyword() ) :: LangExtract.Template.t()
Builds a validated extraction template.
Examples are given as plain maps (string or atom keys, so JSON-loaded
task definitions work verbatim) or as ready-made structs. Map-authored
attributes are normalized to string keys, matching the wire format's
decoded shape; ready-made structs pass through unchanged. Each example's
extraction texts are validated against the example text using the
production aligner; misaligned examples raise
LangExtract.Prompt.Validator.ValidationError — a template that
constructs is a template whose examples align. Pass validate: false
to skip. For runtime data where raising is inappropriate, template/2
returns tagged tuples instead.
Examples
iex> template =
...> LangExtract.template!("Extract conditions.",
...> examples: [
...> %{text: "Patient has diabetes.",
...> extractions: [%{class: "condition", text: "diabetes"}]}
...> ]
...> )
iex> [example] = template.examples
iex> example.extractions
[%LangExtract.Extraction{class: "condition", text: "diabetes", attributes: %{}}]