Unicode.CharacterName (Unicode v2.2.0)

Copy Markdown View Source

Resolves Unicode character names to their codepoint, and back again.

Names are taken from the Name field of the Unicode Character Database (UnicodeData.txt). to_codepoint/1 matches them loosely: case, whitespace, _ and - are ignored (as in \N{...} name lookups). Codepoints whose name is a bracketed label such as <control> have no Name property, so to_name/1 does not resolve them, but the names they are known by are recorded as aliases and to_codepoint/1 does.

Name aliases

NameAliases.txt gives additional names for the characters that need them, because a Name can never change once published: the control characters, which have no Name, and the characters whose published name contains an error. to_codepoint/1 resolves all five alias types, so NULL, LINE FEED, LF, BYTE ORDER MARK and LATIN CAPITAL LETTER GHA each name their character.

Aliases are consulted only after the Name property and the derivation rules, so a name always wins over an alias spelled the same way. aliases/1 returns the aliases of a codepoint with their types.

The names are prefix-compressed (front-coded) into a single sorted binary blob with block restart points, and looked up with a binary search over the restart names followed by a scan-decode within one block. This keeps the table compact without materialising a large map.

to_name/1 is the reverse lookup. Each name is stored once, in normalized form: rather than keeping a second copy of the original, the original is reconstructed by upper-casing and reinserting separators recorded at about 3 bytes per name. That is exact because the Name property contains no lower case. Reverse lookup therefore costs a codepoint index and those separators, not a second name table, and leaves to_codepoint/1 comparing whole binaries as before.

Derived names

CJK ideographs, Tangut ideographs, Hangul syllables and the Seal and Jurchen characters are recorded in UnicodeData.txt as <..., First>/<..., Last> range pairs with no per-character name, and their names are instead derived by the rules in UAX #44. These resolve, but are not part of the name table: they are computed from 15 range tuples and the Hangul jamo short names, under 4KB in total. Materialising them would add more than 131,000 names — over three times the size of the whole listed table — for characters whose names follow a rule.

Derivation is attempted only when the table lookup misses, so it costs nothing on the common path.

Summary

Functions

Returns the Name_Alias values of a codepoint.

Returns the number of names in the table.

Returns the codepoint for a Unicode character name.

Returns the Unicode character name for a codepoint.

Functions

aliases(codepoint)

(since 2.2.0)
@spec aliases(non_neg_integer()) :: [{atom(), String.t()}]

Returns the Name_Alias values of a codepoint.

A character's Name can never change once published, so NameAliases.txt carries the additional names a character needs: the control characters, which have no Name at all, and the characters whose published name contains an error.

Arguments

  • codepoint is a codepoint in the range 0..0x10FFFF.

Returns

  • A list of {type, name} tuples in the order the Unicode Character Database lists them, where type is one of :correction, :control, :alternate, :figment or :abbreviation.

  • An empty list if the codepoint has no aliases.

Examples

iex> Unicode.CharacterName.aliases(0x0000)
[control: "NULL", abbreviation: "NUL"]

iex> Unicode.CharacterName.aliases(0x01A2)
[correction: "LATIN CAPITAL LETTER GHA"]

iex> Unicode.CharacterName.aliases(0xFEFF)
[alternate: "BYTE ORDER MARK", abbreviation: "BOM", abbreviation: "ZWNBSP"]

iex> Unicode.CharacterName.aliases(?A)
[]

count()

@spec count() :: non_neg_integer()

Returns the number of names in the table.

to_codepoint(name, options \\ [])

@spec to_codepoint(String.t(), Keyword.t()) :: {:ok, pos_integer()} | :error

Returns the codepoint for a Unicode character name.

Arguments

  • name is a Unicode character name as a string, matched loosely.

  • options is a keyword list of options.

Options

  • :fuzzy enables approximate matching when the name is not found exactly. The value is either true, meaning use the default Jaro distance of 0.8, or a number between 0.0 and 1.0 giving the minimum distance to accept. The default is false, meaning exact matching only.

Returns

  • {:ok, codepoint} or

  • :error if the name is not known, if a fuzzy search found no single best match, or if the :fuzzy option is not one of the forms above.

Fuzzy matching

A fuzzy search succeeds only when it resolves to one name: the closest name by String.jaro_distance/2 must be at least as close as the threshold and strictly closer than every other name. A tie is treated as unresolved and returns :error, so an ambiguous query never silently picks one of several candidates.

Because the threshold is a floor rather than a filter, it does not need to exclude the many names that are similar to any given query — of the roughly forty thousand names, LATIN SMALL LETTER B is close to LATIN SMALL LETTER A but is not the closest.

Matching is against the listed names only; algorithmically derived names such as CJK UNIFIED IDEOGRAPH-4E00 are not fuzzy-matched, since a near miss on the hexadecimal part would name a different character. Name aliases are matched exactly for the same reason: many are abbreviations of two or three letters, where a single character difference is another abbreviation rather than a typo.

Fuzzy matching scans every name and is several thousand times slower than an exact lookup. It is only attempted after an exact lookup has failed, so supplying the option costs nothing when the name is correct.

Examples

iex> Unicode.CharacterName.to_codepoint("LATIN SMALL LETTER A")
{:ok, 97}

iex> Unicode.CharacterName.to_codepoint("bullet")
{:ok, 8226}

iex> Unicode.CharacterName.to_codepoint("Not A Real Name")
:error

iex> Unicode.CharacterName.to_codepoint("NULL")
{:ok, 0}

iex> Unicode.CharacterName.to_codepoint("LF")
{:ok, 10}

iex> Unicode.CharacterName.to_codepoint("LATIN SMALL LETER A", fuzzy: true)
{:ok, 97}

iex> Unicode.CharacterName.to_codepoint("GRINING FACE", fuzzy: 0.9)
{:ok, 128512}

iex> Unicode.CharacterName.to_codepoint("LATIN SMALL LETER A")
:error

to_name(codepoint)

(since 2.1.0)
@spec to_name(non_neg_integer()) :: {:ok, String.t()} | :error

Returns the Unicode character name for a codepoint.

Arguments

  • codepoint is a codepoint in the range 0..0x10FFFF.

Returns

  • {:ok, name} where name is the Name property of the codepoint, or

  • :error if the codepoint has no name. That includes control characters, surrogates, private use characters and unassigned codepoints, whose Name property is empty. aliases/1 returns the names a control character is known by, which is what to_codepoint/1 resolves.

Notes

Reverse lookup shares the name table with to_codepoint/1 rather than keeping its own copy of every name. It adds a 6 byte per name codepoint index and about 3 bytes per name of separator positions. Names that follow a derivation rule are computed instead of stored and cost nothing.

Where two names differ only by a hyphen that loose matching removes, both resolve here to their own name even though only one of them is reachable through to_codepoint/1.

Examples

iex> Unicode.CharacterName.to_name(0x4E00)
{:ok, "CJK UNIFIED IDEOGRAPH-4E00"}

iex> Unicode.CharacterName.to_name(0xAC00)
{:ok, "HANGUL SYLLABLE GA"}

iex> Unicode.CharacterName.to_name(0x0000)
:error