Text.Extract.Escape (Text v1.0.0)

Copy Markdown View Source

Minimal escaping of URLs, per UTS #58 §4.1.

A URL that percent-escapes every non-ASCII byte is unreadable — https://example.com/%CE%B1%CE%B2 says nothing to a reader, where https://example.com/αβ says αβ. Minimal escaping produces the most readable serialisation that still survives link detection: a character is left as itself unless unescaping it would change where the link ends or what its structure is.

A character therefore stays escaped when it:

  • terminates a link — anything with Link_Term=Hard, such as a space;

  • is the last character of the URL and has Link_Term=Soft, since a trailing soft character is trimmed by the termination algorithm. A soft character anywhere earlier is safe, which is why /αβγ./δεζ. escapes only its very last full stop;

  • would begin a later part — ? and # in a path, # in a query — because unescaping it would silently restructure the URL;

  • is a bracket with no partner in the same field, which would otherwise terminate the link; or

  • is a percent sign followed by two hexadecimal digits, which would otherwise be read as an escape. A % that cannot be misread is left alone.

Everything else is decoded, including all non-ASCII.

This is the counterpart to Text.Extract.Link: that decides where a link ends when reading, this decides how to write one so that reading it back gives the same answer.

Summary

Functions

Rewrites a fully-escaped URL into its minimally escaped, most readable form.

Functions

minimal(url)

@spec minimal(String.t() | keyword()) :: String.t()

Rewrites a fully-escaped URL into its minimally escaped, most readable form.

Arguments

  • url is either a URL string with every non-ASCII byte percent-escaped, or a keyword list of already-parsed parts: :scheme, :host, and any number of :path, :query, :value and :fragment entries in order. Duplicate keys are expected — one :path per segment, and a :value following the :query it belongs to.

The structured form exists because a serialised URL is lossy about its own structure: nothing in https://example.com/α#β distinguishes a path segment containing # from a path followed by a fragment. Pass parts when the structure is already known and the distinction matters.

Returns

  • The URL with escapes removed wherever doing so does not change how the link is detected or structured. The scheme and host are returned unchanged, since host encoding is an IDNA question rather than a link-detection one.

Examples

iex> Text.Extract.Escape.minimal("https://example.com/%CE%B1")
"https://example.com/α"

iex> Text.Extract.Escape.minimal("https://example.com/%CE%B1%20%CE%B2")
"https://example.com/α%20β"

iex> Text.Extract.Escape.minimal("https://example.com/%CE%B4%2E")
"https://example.com/δ%2E"

iex> Text.Extract.Escape.minimal("https://example.com/%CE%B1%3F%CE%BC")
"https://example.com/α%3Fμ"