Resolves Unicode character names to their codepoint, and back again.
Names are taken from the Name field of the Unicode Character Database
(UnicodeData.txt). to_codepoint/1 matches them loosely: case, whitespace,
_ and - are ignored (as in \N{...} name lookups). Codepoints whose name
is a bracketed label such as <control> have no Name property and are not
resolvable in either direction.
The names are prefix-compressed (front-coded) into a single sorted binary blob with block restart points, and looked up with a binary search over the restart names followed by a scan-decode within one block. This keeps the table compact without materialising a large map.
to_name/1 is the reverse lookup. Each name is stored once, in normalized
form: rather than keeping a second copy of the original, the original is
reconstructed by upper-casing and reinserting separators recorded at about 3
bytes per name. That is exact because the Name property contains no lower
case. Reverse lookup therefore costs a codepoint index and those separators,
not a second name table, and leaves to_codepoint/1 comparing whole binaries
as before.
Derived names
CJK ideographs, Tangut ideographs, Hangul syllables and the Seal and Jurchen
characters are recorded in UnicodeData.txt as <..., First>/<..., Last>
range pairs with no per-character name, and their names are instead derived by
the rules in UAX #44. These
resolve, but are not part of the name table: they are computed from 15 range
tuples and the Hangul jamo short names, under 4KB in total. Materialising them
would add more than 131,000 names — over three times the size of the whole
listed table — for characters whose names follow a rule.
Derivation is attempted only when the table lookup misses, so it costs nothing on the common path.
Summary
Functions
Returns the number of names in the table.
Returns the codepoint for a Unicode character name.
Returns the Unicode character name for a codepoint.
Functions
@spec count() :: non_neg_integer()
Returns the number of names in the table.
@spec to_codepoint(String.t(), Keyword.t()) :: {:ok, pos_integer()} | :error
Returns the codepoint for a Unicode character name.
Arguments
nameis a Unicode character name as a string, matched loosely.optionsis a keyword list of options.
Options
:fuzzyenables approximate matching when the name is not found exactly. The value is eithertrue, meaning use the default Jaro distance of0.8, or a number between0.0and1.0giving the minimum distance to accept. The default isfalse, meaning exact matching only.
Returns
{:ok, codepoint}or:errorif the name is not known, if a fuzzy search found no single best match, or if the:fuzzyoption is not one of the forms above.
Fuzzy matching
A fuzzy search succeeds only when it resolves to one name: the closest name by
String.jaro_distance/2 must be at least as close as the threshold and strictly closer than
every other name. A tie is treated as unresolved and returns :error, so an ambiguous query
never silently picks one of several candidates.
Because the threshold is a floor rather than a filter, it does not need to exclude the many names
that are similar to any given query — of the roughly forty thousand names, LATIN SMALL LETTER B
is close to LATIN SMALL LETTER A but is not the closest.
Matching is against the listed names only; algorithmically derived names such as
CJK UNIFIED IDEOGRAPH-4E00 are not fuzzy-matched, since a near miss on the hexadecimal part
would name a different character.
Fuzzy matching scans every name and is several thousand times slower than an exact lookup. It is only attempted after an exact lookup has failed, so supplying the option costs nothing when the name is correct.
Examples
iex> Unicode.CharacterName.to_codepoint("LATIN SMALL LETTER A")
{:ok, 97}
iex> Unicode.CharacterName.to_codepoint("bullet")
{:ok, 8226}
iex> Unicode.CharacterName.to_codepoint("Not A Real Name")
:error
iex> Unicode.CharacterName.to_codepoint("LATIN SMALL LETER A", fuzzy: true)
{:ok, 97}
iex> Unicode.CharacterName.to_codepoint("GRINING FACE", fuzzy: 0.9)
{:ok, 128512}
iex> Unicode.CharacterName.to_codepoint("LATIN SMALL LETER A")
:error
@spec to_name(non_neg_integer()) :: {:ok, String.t()} | :error
Returns the Unicode character name for a codepoint.
Arguments
codepointis a codepoint in the range0..0x10FFFF.
Returns
{:ok, name}wherenameis theNameproperty of the codepoint, or:errorif the codepoint has no name. That includes control characters, surrogates, private use characters and unassigned codepoints, whoseNameproperty is empty.
Notes
Reverse lookup shares the name table with to_codepoint/1 rather than keeping its own copy of
every name. It adds a 6 byte per name codepoint index and about 3 bytes per name of separator
positions. Names that follow a derivation rule are computed instead of stored and cost nothing.
Where two names differ only by a hyphen that loose matching removes, both resolve here to their
own name even though only one of them is reachable through to_codepoint/1.
Examples
iex> Unicode.CharacterName.to_name(0x4E00)
{:ok, "CJK UNIFIED IDEOGRAPH-4E00"}
iex> Unicode.CharacterName.to_name(0xAC00)
{:ok, "HANGUL SYLLABLE GA"}
iex> Unicode.CharacterName.to_name(0x0000)
:error