Minimal escaping of URLs, per UTS #58 §4.1.
A URL that percent-escapes every non-ASCII byte is unreadable — https://example.com/%CE%B1%CE%B2
says nothing to a reader, where https://example.com/αβ says αβ. Minimal escaping produces the
most readable serialisation that still survives link detection: a character is left as itself
unless unescaping it would change where the link ends or what its structure is.
A character therefore stays escaped when it:
terminates a link — anything with
Link_Term=Hard, such as a space;is the last character of the URL and has
Link_Term=Soft, since a trailing soft character is trimmed by the termination algorithm. A soft character anywhere earlier is safe, which is why/αβγ./δεζ.escapes only its very last full stop;would begin a later part —
?and#in a path,#in a query — because unescaping it would silently restructure the URL;is a bracket with no partner in the same field, which would otherwise terminate the link; or
is a percent sign followed by two hexadecimal digits, which would otherwise be read as an escape. A
%that cannot be misread is left alone.
Everything else is decoded, including all non-ASCII.
This is the counterpart to Text.Extract.Link: that decides where a link ends when reading, this
decides how to write one so that reading it back gives the same answer.
Summary
Functions
Rewrites a fully-escaped URL into its minimally escaped, most readable form.
Functions
Rewrites a fully-escaped URL into its minimally escaped, most readable form.
Arguments
urlis either a URL string with every non-ASCII byte percent-escaped, or a keyword list of already-parsed parts::scheme,:host, and any number of:path,:query,:valueand:fragmententries in order. Duplicate keys are expected — one:pathper segment, and a:valuefollowing the:queryit belongs to.
The structured form exists because a serialised URL is lossy about its own structure: nothing in
https://example.com/α#β distinguishes a path segment containing # from a path followed by a
fragment. Pass parts when the structure is already known and the distinction matters.
Returns
- The URL with escapes removed wherever doing so does not change how the link is detected or structured. The scheme and host are returned unchanged, since host encoding is an IDNA question rather than a link-detection one.
Examples
iex> Text.Extract.Escape.minimal("https://example.com/%CE%B1")
"https://example.com/α"
iex> Text.Extract.Escape.minimal("https://example.com/%CE%B1%20%CE%B2")
"https://example.com/α%20β"
iex> Text.Extract.Escape.minimal("https://example.com/%CE%B4%2E")
"https://example.com/δ%2E"
iex> Text.Extract.Escape.minimal("https://example.com/%CE%B1%3F%CE%BC")
"https://example.com/α%3Fμ"