Text.Extract.Link (Text v1.0.0)

Copy Markdown View Source

Phase 3 of the URL / email extraction pipeline: decide where a candidate link ends.

This implements the termination algorithm of UTS #58 §3.5.1, replacing the ASCII-only heuristics that preceded it. Real-world prose embeds links in sentences, so See http://example.com. must yield http://example.com with the sentence-final full stop dropped, while https://en.wikipedia.org/wiki/URI_(disambiguation) must keep its closing parenthesis because a matching opener appears inside the link.

UTS #58 does not define where a link starts — §3.2 puts that outside its scope — so Text.Extract.Scanner still finds candidates and Text.Extract.Url still validates them. Only the end of the span is decided here.

How it works

The scan walks the candidate one codepoint at a time, keeping a last_safe offset that marks the longest prefix known to be a valid link. Each codepoint is classified by Unicode.LinkTerm:

  • :include — part of the link; last_safe advances past it.

  • :hard — ends the link immediately; the scan returns last_safe.

  • :soft — provisionally part of the link, but last_safe does not advance. A later :include pulls it back in, so a.b keeps its full stop while a. does not.

  • :open — a bracket, pushed onto a stack. last_safe does not advance, so a trailing unclosed bracket is dropped.

  • :close — pops the stack and compares through Unicode.LinkBracket. A match is included and last_safe advances; a mismatch, or a close with nothing on the stack, ends the link.

The bracket stack is cleared at separators, so a bracket opened in one field cannot be closed in the next: at the part initiators /, ? and # always, at = and & within a query or fragment, and at a comma within a fragment. Those are scoped to their part — = is an ordinary path character, and clearing on it everywhere would break URLs such as .../system.net.httpwebrequest(v=VS.100).aspx. The stack is capped at 125 entries, beyond which further open brackets are treated as ordinary included characters rather than being stacked.

Because the properties cover the whole repertoire rather than a handful of ASCII characters, this handles the 61 non-ASCII bracket pairs and 129 ranges of soft terminators that the previous implementation could not.

Summary

Functions

Trims a candidate link to the longest prefix that UTS #58 considers part of the link.

Functions

shrink(candidate)

@spec shrink(String.t()) :: String.t()

Trims a candidate link to the longest prefix that UTS #58 considers part of the link.

Arguments

Returns

  • The candidate with any trailing characters that terminate the link removed. Never grows; only the end of the string is trimmed.

Examples

iex> Text.Extract.Link.shrink("http://example.com.")
"http://example.com"

iex> Text.Extract.Link.shrink("http://example.com)")
"http://example.com"

iex> Text.Extract.Link.shrink("http://en.wikipedia.org/wiki/URI_(disambiguation)")
"http://en.wikipedia.org/wiki/URI_(disambiguation)"

iex> Text.Extract.Link.shrink("http://x.com/path......")
"http://x.com/path"