Phase 3 of the URL / email extraction pipeline: decide where a candidate link ends.
This implements the termination algorithm of
UTS #58 §3.5.1, replacing the ASCII-only heuristics that
preceded it. Real-world prose embeds links in sentences, so See http://example.com. must yield
http://example.com with the sentence-final full stop dropped, while
https://en.wikipedia.org/wiki/URI_(disambiguation) must keep its closing parenthesis because a
matching opener appears inside the link.
UTS #58 does not define where a link starts — §3.2 puts that outside its scope — so
Text.Extract.Scanner still finds candidates and Text.Extract.Url still validates them. Only
the end of the span is decided here.
How it works
The scan walks the candidate one codepoint at a time, keeping a last_safe offset that marks the
longest prefix known to be a valid link. Each codepoint is classified by Unicode.LinkTerm:
:include— part of the link;last_safeadvances past it.:hard— ends the link immediately; the scan returnslast_safe.:soft— provisionally part of the link, butlast_safedoes not advance. A later:includepulls it back in, soa.bkeeps its full stop whilea.does not.:open— a bracket, pushed onto a stack.last_safedoes not advance, so a trailing unclosed bracket is dropped.:close— pops the stack and compares throughUnicode.LinkBracket. A match is included andlast_safeadvances; a mismatch, or a close with nothing on the stack, ends the link.
The bracket stack is cleared at separators, so a bracket opened in one field cannot be closed in
the next: at the part initiators /, ? and # always, at = and & within a query or
fragment, and at a comma within a fragment. Those are scoped to their part — = is an ordinary
path character, and clearing on it everywhere would break URLs such as
.../system.net.httpwebrequest(v=VS.100).aspx. The stack is capped at 125 entries, beyond which
further open brackets are treated as ordinary included characters rather than being stacked.
Because the properties cover the whole repertoire rather than a handful of ASCII characters, this handles the 61 non-ASCII bracket pairs and 129 ranges of soft terminators that the previous implementation could not.
Summary
Functions
Trims a candidate link to the longest prefix that UTS #58 considers part of the link.
Functions
Trims a candidate link to the longest prefix that UTS #58 considers part of the link.
Arguments
candidateis the candidate substring, as emitted byText.Extract.Scanner.scan/1.
Returns
- The candidate with any trailing characters that terminate the link removed. Never grows; only the end of the string is trimmed.
Examples
iex> Text.Extract.Link.shrink("http://example.com.")
"http://example.com"
iex> Text.Extract.Link.shrink("http://example.com)")
"http://example.com"
iex> Text.Extract.Link.shrink("http://en.wikipedia.org/wiki/URI_(disambiguation)")
"http://en.wikipedia.org/wiki/URI_(disambiguation)"
iex> Text.Extract.Link.shrink("http://x.com/path......")
"http://x.com/path"