All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
Unreleased
0.6.0 - 2026-07-05
Security
- Providers no longer follow HTTP redirects — Req strips only the
standard
authorizationheader on cross-host redirects, so Claude'sx-api-keywould have been forwarded to a redirect target. LLM APIs never legitimately redirect these POSTs; a 3xx now surfaces as{:error, {:api_error, status, body}}. Re-enable viareq_options: [redirect: true]if you proxy through something that redirects.
Changed
- Prompts adopt upstream's Q/A scaffold —
Examplesheading,Q:/A:pairs, and a trailing bareA:answer primer, mirroring langextract'sQAPromptGenerator. Measured effect: ~20% fewer output tokens (reduced adaptive-thinking spend), no alignment cost. reqconstraint tightened to~> 0.6— the previous~> 0.5admitted pre-1.0 minors the test suite has never run against.WireFormat.normalize/1parses JSON first — the strict, fast parser handles the (now default) JSON responses; the YAML parser and its repair pass remain as the tolerance path for models that answer in YAML.- Wire format is now JSON (was YAML) —
WireFormat.format_extractions/1emits fenced dynamic-key JSON, matching upstream's default; theymlrdependency is dropped. Decided by a corpus A/B under the new scaffold: zero chunk errors across 440 quote-dense dialogue chunks, better dialogue alignment (17 fuzzy / 0 not_found vs 68 / 2), and 25% fewer ner output tokens. Decoding is format-agnostic —WireFormat.normalize/1accepts JSON and YAML responses alike and keeps the YAML repair machinery.
Fixed
Serializer.from_map/1validates extraction entries — entries missing"text"or carrying a non-string"class"now return the promised{:error, :invalid_data}instead of producing malformed spans. Class-less spans (fromalign/3) still round-trip.- Chunk task timeouts now return the documented error tuple — the
{:error, {:task_exit, reason}}shape promised byrun/4was unreachable:Task.async_stream's defaulton_timeout: :exitcrashed the calling process instead.on_timeout: :kill_taskmakes a timed-out chunk surface as the documented infrastructure-failure return.
Added
- Telemetry —
[:lang_extract, :request],[:lang_extract, :chunk], and[:lang_extract, :document]spans (:telemetryis now an explicit dependency). Request:stopevents carry input/output token counts normalized across all three providers; the benchmark runners record them as per-documentusageblocks with per-request latency.
0.5.0 - 2026-07-05
Added
- Alignment tuning options —
:min_density(LCS token-density floor, default 1/3),:accept_lesser(toggle prefix matching), and:exact_algorithm(:dp|:first_occurrence), accepted byLangExtract.run/4,extract/3, andalign/3. - "Alignment and Spans" hexdocs guide — span semantics, byte-vs-character
offsets (
binary_part, notString.slice), the four aligner phases, and tuning. The hexdocs sidebar now groups modules by layer, and thealign/extractexamples run as doctests.
Changed
- Aligner ported to upstream langextract v1.6.0 semantics — the fuzzy
phase's frequency-overlap sliding window is replaced by upstream's
difflib-style lesser prefix match plus an LCS dynamic program over
lightly stemmed tokens, gated by coverage (
:fuzzy_threshold) and density (:min_density). Behavior is pinned by differential fixtures generated from the upstream aligner. On the dialogue benchmark this took the exact rate from 77.5% to 93.7% — identical to Python's 93.7% on the same corpus. - Every prompt now demands verbatim spans —
Prompt.Builderappends a standing instruction requiring extractions to be exact source substrings and an emptyextractions: []on contentless passages. Benchmarked: without it, dialogue runs produced dozens of few-shot echoes and stitched paraphrases that could not be aligned. - Repeated mentions now ground to successive occurrences — the aligner
gained a phase-0 monotonic occurrence DP (port of upstream #485): over the
extraction list in model output order, it selects at most one exact
occurrence per extraction, order-preserving and non-overlapping,
maximizing matched tokens. Previously every within-chunk repeat took the
first occurrence's offsets — 32% of exact spans in the ner benchmark
landed on an already-claimed position; after the port, offset agreement
with upstream on repeated mentions with matched counts is 90.9%, on par
with unique mentions.
exact_algorithm: :first_occurrencerestores the old behavior. - Chunk stream no longer blocks on the slowest chunk — the orchestrator's
Task.async_streamnow runsordered: false(document order is restored by sorting chunk results on their byte offsets), so one slow chunk — e.g. a 429 riding Req's retry backoff — no longer gates every later chunk launch. The default:max_concurrencyalso rises from 3 to 10, matching upstream langextract'smax_workersdefault. - Minimum Elixir raised to 1.15 — plug 1.20 (test dependency, pulled in by a security patch) requires Elixir 1.15, and CI can no longer verify 1.14.
Fixed
- Claude provider default model updated to
claude-sonnet-5— the previous default,claude-sonnet-4-20250514, was retired upstream on 2026-06-15, soLangExtract.new(:claude)without an explicit:modelreturned 404s. - Gemini provider default model updated to
gemini-3.5-flash— tracking upstream langextract's default (their #472);gemini-2.0-flashis approaching retirement, the same failure class as the Claude default. - Claude provider no longer sends
temperatureby default — claude-sonnet-5 rejects non-default sampling parameters with a 400, so the oldtemperature: 0default broke every request. It is now sent only when the caller explicitly sets:temperature. WireFormat.normalize/1no longer corrupts YAML block scalars — the colon-quoting pass treated block scalar headers (dialogue: |-) as values and quoted them, orphaning the indented lines and failing the parse. Claude Sonnet 5 emits multi-line extractions as block scalars, so this caused chunk-level{:invalid_format, _}failures.WireFormat.normalize/1parses first, repairs only on failure — valid YAML (including multi-line plain scalars) is never rewritten. The repair pass now also recovers unterminated and mis-escaped quoted values and folds plain-scalar continuation lines, fixing all chunk failures observed in the July 2026 benchmark (11/11 payloads, 90 extractions recovered).
0.4.0 - 2026-07-02
Changed
LangExtract.IOrenamed toLangExtract.Serializer(breaking) — The old name shadowed Elixir's standard-libraryIOmodule, forcing callers to alias around the collision. The functions are unchanged.LangExtract.Pipeline.Extractionpromoted toLangExtract.Extraction(breaking) — The struct users build in every template example is the library's central payload, shared byPromptandPipelinealike; it now lives at the top level instead of inside one consumer's namespace.LangExtract.Pipeline.FormatHandlerrenamed toLangExtract.WireFormat(breaking) — The LLM wire-format port (encode for prompts, decode for responses) moved to the top level for the same reason. With both moves,Promptno longer depends onPipelineat all.- Provider HTTP defaults: 120s receive timeout and transient retries —
Reverses the 0.2.0 "retries disabled by default" decision. LLM completions
routinely exceed Req's 15s
receive_timeoutdefault, andretry: falsemeant a transient 429/5xx permanently dropped a chunk as aChunkError. Both remain overridable viareq_options:. - Aligner ports upstream langextract v1.6.0 semantics — After the exact
phase, a difflib-style lesser phase grounds partial matches anchored at the
extraction's first token, and an LCS subsequence fallback (with upstream's
0.75 coverage and 1/3 density gates, plus light plural stemming) replaces
the fixed-window fuzzy matcher. Extractions that previously returned
:not_found(interrupted dialogue, plural variants) now ground as:fuzzywith trimmed spans. New options::min_density,:accept_lesser. Verified against upstream via generated differential fixtures (test/fixtures/alignment_parity.json). - Exact alignment via linear scan — Replaces
List.myers_difference/2, which did O(N²) work in source token count and missed genuinely contiguous matches when extraction tokens also appeared scattered earlier in the source (those fell back to:fuzzy; they now align as:exactwith the same byte offsets). - Hex package no longer ships the benchmark Mix task — the
benchmark.runtask needs the localbenchmark/corpus, which was never packaged, so the task could only fail for downstream users. An explicitfiles:list now scopes the package to the library itself.
Added
- Verbatim extraction instruction in prompts —
Prompt.Buildernow instructs the model to extract only verbatim spans and to emitextractions: []for contentless passages. Reduces ungrounded extractions (few-shot echoes, merged interrupted quotes) that could never align. Serializer.span_to_map/1— Public single-span serialization (previously private), also used by the benchmark task instead of a duplicated implementation.
Fixed
Serializer.from_map/1andload_jsonl/1no longer raise on malformed input — An unknown or missing extraction"status"now returns{:error, :invalid_data}(the module's existing error contract) instead of raisingArgumentErrorfromString.to_existing_atom/1.
0.3.0 - 2026-04-06
Changed
LangExtract.run/4returns{:ok, {spans, chunk_errors}} | {:error, reason}— Always returns partial results alongside chunk errors instead of halting on the first failure. Infrastructure failures (task exits, timeouts) return{:error, reason}.- Pipeline namespace —
FormatHandler,Parser,Extractionmoved underLangExtract.Pipeline.*.Pipelineis the public API for the extraction context. - YAML format with quoting — LLM wire format switched from JSON to YAML, matching the upstream Python library. Unquoted values containing colons are automatically quoted before parsing.
- Removed
run_single— All text goes through chunking, matching Python's behavior. Themax_chunk_chars: :disabledoption is removed. - Removed
:on_chunk_errorcallback — Errors are now visible in the return value. The callback was redundant. - Removed previous chunk context — Was causing cross-chunk
not_foundalignments. Python disables this by default. Chunkstruct now includesbyte_end, computed once inpack_sentences.FormatHandler.normalize/1passes through valid YAML without anextractionskey, lettingParserreturn:missing_extractions.- Broke dependency cycle between
LangExtractandOrchestrator. Shared pipeline logic (normalize → parse → align) extracted intoLangExtract.Pipeline. - Reuse Req HTTP client across requests. New
build_http_client/1callback onProviderbehaviour builds theReqstruct once atnew/2time, stored onClient.http_clientand reused for all subsequent requests. - Tokenizer
classify/1uses binary pattern matching for ASCII bytes, falling back to Unicode regex only for non-ASCII. Avoids up to 3 regex calls per token. Clientstruct now redacts:optionsand:http_clientfrominspectoutput to prevent accidental API key exposure in logs.
Added
LangExtract.Pipeline.ChunkError— Struct withbyte_start,byte_end, andreasonfor failed chunk regions.- Benchmark improvements — Per-document JSON files in timestamped
directories,
--documentflag for single-document runs,_latestsymlink.
0.2.2 - 2026-03-19
Added
- ROADMAP.md — Documents future improvements and unported features from the original Python library.
- Aligner edge-case tests — Additional test coverage inspired by the Python langextract test suite.
Changed
- README.md — Moved future improvements to ROADMAP.md. Cleaned up comparison section.
Removed
docs/directory — Removed historical design specs and implementation plans (17 files, ~7,000 lines). These served their purpose during development; the project is now documented via README, CHANGELOG, and ROADMAP.
0.2.1 - 2026-03-19
Fixed
- Remove stale
httpowerentry frommix.lock.
0.2.0 - 2026-03-19
Changed
- Replaced HTTPower with Req as the HTTP client. Req is a mature,
batteries-included HTTP client with wide ecosystem adoption. This removes
the
httpowerand directfinchdependencies. - Gemini API key now passed via Req's
params:option instead of being embedded in the URL path string. - Req retries disabled by default in all providers. Callers can opt in
via
req_options: [retry: :transient]. - Generic
:req_optionspassthrough replaces the test-specific:plugoption. Any Req configuration (timeouts, retry, pool settings, plug for testing) can be forwarded to the underlying Req request.
Added
- Orchestrator with chunking —
LangExtract.run/3,4wires the full pipeline end-to-end. Sentence-aware chunking via:max_chunk_charsoption withTask.async_streamfor parallel inference. LangExtract.new/2— Req-inspired two-step API: create a client, then run extractions.LangExtract.Chunker— Sentence-aware text splitting with abbreviation awareness and three-tier strategy.LangExtract.IO— Serialize extraction results to plain maps and JSONL.- Module reorganization — Alignment and Prompt subdomains for cleaner namespace organization.
0.1.0 - 2026-03-18
Initial release. A complete Elixir port of the core pipeline from google/langextract — extracts structured data from text using LLMs and maps every extraction back to exact byte positions in the source.
Added
Core Pipeline
LangExtract.new/2— Create a configured LLM client with a provider shorthand (:claude,:openai,:gemini) and provider-specific options.LangExtract.run/3,4— Run the full extraction pipeline: build prompt → call LLM → normalize → parse → align → return enriched spans.LangExtract.extract/3— Parse raw LLM output and align extractions against source text. Accepts both canonical (class/text/attributes) and dynamic-key format.LangExtract.align/3— Align extraction strings to byte spans in source text without LLM involvement.
Alignment (LangExtract.Alignment.*)
- Tokenizer — Regex-based tokenizer producing tokens with byte offsets. Keeps contractions as single tokens for better English alignment.
- Two-phase Aligner — Phase 1: exact contiguous match via
List.myers_difference/2. Phase 2: fuzzy sliding-window fallback with configurable threshold (default 0.75). Uses tuples for O(1) index access. - Span struct — Holds extraction text, byte offsets (
byte_start,byte_end), alignment status (:exact,:fuzzy,:not_found), plus optionalclassandattributesfrom the LLM.
Prompt Building (LangExtract.Prompt.*)
- Template — Struct holding a task description and few-shot examples.
- ExampleData — Struct for a single few-shot example (source text + expected extractions).
- Builder — Renders Q&A-formatted prompts with dynamic-key extraction
examples. Supports cross-chunk context via
:previous_chunkoption. - Validator — Pre-flight check that few-shot examples align against their
own source text.
validate/1returns results;validate!/1raises. The caller decides severity — no built-in logging or severity levels.
Format Handler
LangExtract.Pipeline.FormatHandler— Hexagonal port between external LLM format and internal domain. SerializesExtractionstructs to dynamic-key JSON for prompts. Normalizes raw LLM output (strips<think>tags, markdown fences, converts dynamic keys to canonicalclass/text/attributesformat). Returns decoded maps to avoid redundant JSON round-trips.
LLM Providers
- Provider behaviour — Single
infer/2callback. Shared helpers for API key resolution (fetch_api_key/2), common options (common_opts/2), and HTTP error mapping (map_response/2). - Claude (
LangExtract.Provider.Claude) — Anthropic Messages API via Req.x-api-keyheader auth. - OpenAI (
LangExtract.Provider.OpenAI) — Chat Completions API via Req. Bearer auth. Optional JSON mode (:json_modeoption, defaulttrue). Works with any OpenAI-compatible endpoint. - Gemini (
LangExtract.Provider.Gemini) — REST API via Req. Query parameter auth. JSON output viaresponseMimeType.
Chunking
LangExtract.Chunker— Sentence-aware text chunking with three-tier strategy: sentence packing → newline splitting → token fallback. Abbreviation-aware sentence detection (Mr.,Dr., etc.). Newline + uppercase heuristic for paragraph breaks.- Orchestrator chunking — When
:max_chunk_charsis set, the orchestrator splits the source, processes chunks in parallel viaTask.async_stream, adjusts byte offsets, and concatenates results. Previous chunk text is passed as prompt context for cross-chunk coreference resolution.
I/O
LangExtract.IO— Serialize extraction results to plain maps (to_map/2) and back (from_map/1). Save/load multiple results as JSONL (save_jsonl/2,load_jsonl/1).
Infrastructure
- Client struct — Holds provider module and options. Created via
LangExtract.new/2. - Req — Batteries-included HTTP client. Uses
json:option for automatic request body encoding. Retries disabled by default; opt in via:req_options. - Req.Test — All provider integration tests use stubs, not network calls.
- Credo — Strict mode passes with zero issues.
- 187 tests — Full coverage across all modules.
Divergences from Python Reference
- Byte offsets instead of character offsets (natural for Elixir binaries).
- Contraction handling —
don'tis one token, not three. - No
MATCH_LESSER/MATCH_GREATER— Deliberate simplification. Our three statuses (:exact,:fuzzy,:not_found) are cleaner. - Claude provider — Not in the original; added as the primary provider.
- JSON only — No YAML support (modern LLMs handle JSON well).
- Caller-decides severity for prompt validation (no built-in severity enum).
- Req-inspired API —
new/2+run/3,4instead of a single function with many keyword arguments.