Unicode.CharacterName (Unicode v2.1.0)

Copy Markdown View Source

Resolves Unicode character names to their codepoint, and back again.

Names are taken from the Name field of the Unicode Character Database (UnicodeData.txt). to_codepoint/1 matches them loosely: case, whitespace, _ and - are ignored (as in \N{...} name lookups). Codepoints whose name is a bracketed label such as <control> have no Name property and are not resolvable in either direction.

The names are prefix-compressed (front-coded) into a single sorted binary blob with block restart points, and looked up with a binary search over the restart names followed by a scan-decode within one block. This keeps the table compact without materialising a large map.

to_name/1 is the reverse lookup. Each name is stored once, in normalized form: rather than keeping a second copy of the original, the original is reconstructed by upper-casing and reinserting separators recorded at about 3 bytes per name. That is exact because the Name property contains no lower case. Reverse lookup therefore costs a codepoint index and those separators, not a second name table, and leaves to_codepoint/1 comparing whole binaries as before.

Derived names

CJK ideographs, Tangut ideographs, Hangul syllables and the Seal and Jurchen characters are recorded in UnicodeData.txt as <..., First>/<..., Last> range pairs with no per-character name, and their names are instead derived by the rules in UAX #44. These resolve, but are not part of the name table: they are computed from 15 range tuples and the Hangul jamo short names, under 4KB in total. Materialising them would add more than 131,000 names — over three times the size of the whole listed table — for characters whose names follow a rule.

Derivation is attempted only when the table lookup misses, so it costs nothing on the common path.

Summary

Functions

Returns the number of names in the table.

Returns the codepoint for a Unicode character name.

Returns the Unicode character name for a codepoint.

Functions

count()

@spec count() :: non_neg_integer()

Returns the number of names in the table.

to_codepoint(name, options \\ [])

@spec to_codepoint(String.t(), Keyword.t()) :: {:ok, pos_integer()} | :error

Returns the codepoint for a Unicode character name.

Arguments

  • name is a Unicode character name as a string, matched loosely.

  • options is a keyword list of options.

Options

  • :fuzzy enables approximate matching when the name is not found exactly. The value is either true, meaning use the default Jaro distance of 0.8, or a number between 0.0 and 1.0 giving the minimum distance to accept. The default is false, meaning exact matching only.

Returns

  • {:ok, codepoint} or

  • :error if the name is not known, if a fuzzy search found no single best match, or if the :fuzzy option is not one of the forms above.

Fuzzy matching

A fuzzy search succeeds only when it resolves to one name: the closest name by String.jaro_distance/2 must be at least as close as the threshold and strictly closer than every other name. A tie is treated as unresolved and returns :error, so an ambiguous query never silently picks one of several candidates.

Because the threshold is a floor rather than a filter, it does not need to exclude the many names that are similar to any given query — of the roughly forty thousand names, LATIN SMALL LETTER B is close to LATIN SMALL LETTER A but is not the closest.

Matching is against the listed names only; algorithmically derived names such as CJK UNIFIED IDEOGRAPH-4E00 are not fuzzy-matched, since a near miss on the hexadecimal part would name a different character.

Fuzzy matching scans every name and is several thousand times slower than an exact lookup. It is only attempted after an exact lookup has failed, so supplying the option costs nothing when the name is correct.

Examples

iex> Unicode.CharacterName.to_codepoint("LATIN SMALL LETTER A")
{:ok, 97}

iex> Unicode.CharacterName.to_codepoint("bullet")
{:ok, 8226}

iex> Unicode.CharacterName.to_codepoint("Not A Real Name")
:error

iex> Unicode.CharacterName.to_codepoint("LATIN SMALL LETER A", fuzzy: true)
{:ok, 97}

iex> Unicode.CharacterName.to_codepoint("GRINING FACE", fuzzy: 0.9)
{:ok, 128512}

iex> Unicode.CharacterName.to_codepoint("LATIN SMALL LETER A")
:error

to_name(codepoint)

(since 2.1.0)
@spec to_name(non_neg_integer()) :: {:ok, String.t()} | :error

Returns the Unicode character name for a codepoint.

Arguments

  • codepoint is a codepoint in the range 0..0x10FFFF.

Returns

  • {:ok, name} where name is the Name property of the codepoint, or

  • :error if the codepoint has no name. That includes control characters, surrogates, private use characters and unassigned codepoints, whose Name property is empty.

Notes

Reverse lookup shares the name table with to_codepoint/1 rather than keeping its own copy of every name. It adds a 6 byte per name codepoint index and about 3 bytes per name of separator positions. Names that follow a derivation rule are computed instead of stored and cost nothing.

Where two names differ only by a hyphen that loose matching removes, both resolve here to their own name even though only one of them is reachable through to_codepoint/1.

Examples

iex> Unicode.CharacterName.to_name(0x4E00)
{:ok, "CJK UNIFIED IDEOGRAPH-4E00"}

iex> Unicode.CharacterName.to_name(0xAC00)
{:ok, "HANGUL SYLLABLE GA"}

iex> Unicode.CharacterName.to_name(0x0000)
:error