Unicode v2.2.0

This is the changelog for Unicode v2.2.0 released on September 20th, 2026. For older changelogs please consult the release tag on GitHub

This release makes every property value the Unicode Character Database declares resolvable, including the values a data file supplies only through an @missing annotation.

Enhancements

  • Resolves the default value of fifteen further properties, so \p{jt=U}, \p{bpt=None}, \p{sc=Unknown}, \p{age=NA} and their companions now answer. Every property's values together account for all 1,114,112 codepoints, where before only five did.

  • Applies the range-specific defaults BidiClass.txt declares, so an unassigned codepoint in a right-to-left block resolves as R or AL rather than L.

  • Resolves Joining_Group values from any spelling, including the Crown_Ain form PropertyValueAliases.txt uses. 73 of the 116 values previously resolved from neither spelling.

  • Resolves binary property aliases that contain a separator, such as Bidi_M, Gr_Base and Pat_Syn, which the lookup normalised but the alias table did not.

  • Documents Unicode.fetch_property/1 and Unicode.get_property/1, which resolve a property named at runtime to the module that serves it. Both were public but carried @doc false, so the lookup that unicode_set evaluates set expressions with was absent from the documentation.

  • Adds an introduction guide covering the shape of the API: the three ways to reach a property, naming property values, character names, script extensions and the guards. Its examples run as doctests.

  • Loads NameAliases.txt, so Unicode.CharacterName.to_codepoint/2 resolves the names of the control characters, which have no Name property, along with abbreviations such as LF and the corrections for characters whose published name contains an error. Unicode.CharacterName.aliases/1 returns a codepoint's aliases with their types.

Bug fixes

Notes

  • known_*/0 and count/1 include the default value for the affected properties, so those lists and counts are larger than in 2.1.0.

Unicode v2.1.0

This release closes a number of Unicode Character Database coverage gaps: several enumerated properties that appeared in PropertyAliases.txt but had no backing data are now resolvable through Unicode.fetch_property/1 and have their own introspection modules.

Enhancements

  • Resolves the property value that a UCD @missing annotation declares as the default. Grapheme_Cluster_Break, Word_Break and Sentence_Break now answer Other (and its XX alias), and Line_Break answers XX, for every codepoint the data files do not list. Line_Break previously covered only 358,578 of the 1,114,112 codepoints; the remainder, including all private use, resolved only through a lookup fallback and was invisible to anything enumerating the property. Where a value is both listed explicitly and named as the @missing default, as XX is for Line_Break, the two sets are merged rather than replaced.

  • Adds the age property value aliases from PropertyValueAliases.txt, so Unicode.Age.fetch("V18_0") resolves the same set as Unicode.Age.fetch("18.0").

  • Updates the underlying data to the final Unicode 18.0 release (16 September 2026), adding the Jurchen, Proto_Cuneiform and Seal scripts, seven new blocks and ten new Crown_* joining groups.

  • Adds Unicode.ScriptExtensions for the Script_Extensions (scx) property, the last UCD enumerated property without a backing module and a UTS #18 RL1.2 conformance requirement. Unicode.ScriptExtensions.script_extensions/1 returns the set of scripts a codepoint is used with, defaulting to its Script value where the UCD lists no explicit set.

  • Unicode.CharacterName.to_codepoint/2 accepts a :fuzzy option that resolves a misspelled name by String.jaro_distance/2, either as true for the default distance of 0.8 or as an explicit distance. It succeeds only where one name is strictly closest than all others, so an ambiguous query returns :error rather than guessing.

  • Adds Unicode.CharacterName.to_name/1, the reverse of to_codepoint/1, returning the Name property of a codepoint. Each name is still stored only once — the original is reconstructed from the normalized form plus separator positions — so reverse lookup adds about 410KB rather than a second name table, and leaves to_codepoint/1 unchanged in speed.

  • Unicode.CharacterName.to_codepoint/1 now resolves algorithmically derived names — CJK and Tangut ideographs, Hangul syllables, and the Seal and Jurchen characters — covering a further 131,576 characters. The names are computed from the UAX #44 derivation rules on a table-lookup miss rather than stored, adding under 4KB.

  • mix unicode.download now resolves the UCD and emoji trees through independent release channels, settable per run with --release / --emoji-release, by environment variable, or in application config. Adds --into to download to a scratch directory, --dry-run, and a post-download check that every data file reports the same Unicode version.

  • Adds the enumerated property modules Unicode.Age, Unicode.NumericType, Unicode.NumericValue, Unicode.DecompositionType, Unicode.HangulSyllableType, Unicode.IndicPositionalCategory, Unicode.VerticalOrientation, Unicode.JoiningGroup and Unicode.BidiPairedBracketType, each with a codepoint lookup, range introspection and a top-level delegate on Unicode.

  • Adds the normalization quick check modules Unicode.NfcQuickCheck, Unicode.NfdQuickCheck, Unicode.NfkcQuickCheck and Unicode.NfkdQuickCheck (the NFC_QC, NFD_QC, NFKC_QC and NFKD_QC properties) with Unicode.nfc_quick_check/1 and friends.

  • Adds Bidi_Mirrored as a first class boolean property, so Unicode.Property.bidi_mirrored?/1 and Unicode.properties/1 now report it.

  • The Name property now resolves through Unicode.fetch_property/1 to Unicode.CharacterName.

  • Adds the UTS #58 link properties Unicode.LinkTerm, Unicode.LinkEmail and Unicode.LinkBracket, which together drive link detection in flowing text. mix unicode.download gains a linkification tree alongside the UCD and emoji trees to fetch them.

  • Adds the LC (Cased_Letter) general category, the union of Lu, Ll and Lt. It is the only General_Category group whose name is not a single letter, so it was not produced by the derivation that yields L, N, P and the rest, and Unicode.GeneralCategory.fetch("lc") returned :error despite the alias being advertised.

Bug fixes

  • Property aliases containing an underscore (for example nfc_qc) are now reachable through Unicode.fetch_property/1, which previously only matched the whitespace-stripped canonical form.

  • Scripts and blocks whose PropertyValueAliases line carries more than two names now resolve by every spelling. Unicode.Script.fetch/1 failed for "copt" and "qaac", and Unicode.Block.fetch/1 for "latin1sup" and "cyrillicsupplementary", because the alias map was built by inverting a many-to-one map and elected a name absent from the data.

  • Unicode.version/0 no longer reports a stale version after the data files are updated. It derives the version from blocks.txt at compile time but did not declare it as an @external_resource, so the module was not recompiled when the data changed.

  • mix unicode.download no longer requests a non-existent emoji URL. The emoji files are versioned separately from the UCD and Public/emoji/17.0/ was never published, so the previous version-interpolated path returned 404 and the files had to be updated by hand.

Unicode v2.0.0

This is the changelog for Unicode v2.0.0 released on July 9th, 2026. For older changelogs please consult the release tag on GitHub

This is a major release that folds the unicode_guards library into unicode and replaces the hand-maintained derived category tables with values computed directly from the character database.

Breaking changes

  • The Unicode.Guards module is now provided by this library. The separate unicode_guards package is no longer required; remove it from your dependencies and depend on unicode instead.

  • The derived general categories :Assigned, :Graph, :Visible and :Printable now reflect the current Unicode data. The previous static tables were stale: :Assigned gained 15,942 codepoints (283,440 to 299,382), :Graph/:Visible gained the same 15,942 (281,308 to 297,250), and :Printable is corrected to match String.printable?/1 (it previously excluded most of the BMP, including Arabic, CJK and Hangul).

  • The internal Unicode.DerivedCategory.Assigned and Unicode.DerivedCategory.Graph modules have been removed; these categories are now computed at compile time.

Enhancements

  • Adds Unicode.CharacterName.to_codepoint/1, which resolves a Unicode character name to its codepoint with loose matching. The name table is stored as a sorted binary blob with a binary search to keep it compact.

  • Codepoint lookup functions such as Unicode.GeneralCategory.category/1, Unicode.Script.script/1 and the Unicode.Property boolean functions are now implemented with binary search over compact range tables instead of very large generated guard clauses. Compilation is an order of magnitude faster and codepoint lookups are approximately 10x faster.

  • Adds Unicode.RangeSearch which builds the range search tables at compile time and performs the binary search over them.

  • Documentation for all public modules and functions now follows a standard format with arguments, return values and examples.

  • Removes the unused and incorrect Unicode.Utils.remove_reserved_codepoints/1.

  • Adds Unicode.Guards, a set of guards (is_upper/1, is_lower/1, is_digit/1, is_whitespace/1, is_graph/1, the quotation-mark guards and more) for use in function when clauses. Folded in from unicode_guards with no runtime dependencies.

  • Derived categories are computed from the character database using new range-set helpers (Unicode.Utils.union_ranges/1, complement_ranges/1 and difference_ranges/2), removing the need to regenerate static tables by hand each release.

Bug Fixes

  • Fix UTF-16 and UTF-32 validation. Unicode.replace_invalid/3 for the :utf16, :utf16be, :utf16le, :utf32, :utf32be and :utf32le encodings previously crashed on any input and dropped valid codepoints.

  • Fix :utf32 validation dispatching to the UTF-16 implementation.

  • Replacement strings for UTF-16 and UTF-32 validation are now transcoded to the target encoding rather than being spliced in as UTF-8 bytes.

  • Fix Unicode.compact_ranges/1 truncating a range that fully contains the following range.

  • Fix Unicode.Emoji.emoji/1 returning nil for single-codepoint emoji graphemes in a string.

  • Fix Unicode.Block.fetch/1 and Unicode.Block.get/1 for block names containing digits, such as "Latin-1 Supplement" and "Number Forms", which previously returned :error/nil for every spelling of the canonical name because no alias mapped the normalised name to the block key.

Unicode v1.22.0

This is the changelog for Unicode v1.22.0 released on May 4th, 2026. For older changelogs please consult the release tag on GitHub

Enhancements

  • Adds the Bidi_Class Unicode property.

  • Adds the Joining_Type Unicode property.

Unicode v1.21.2

This is the changelog for Unicode v1.21.2 released on April 29th, 2026. For older changelogs please consult the release tag on GitHub

Breaking changes

  • Supported on Elixir 1.17 and later only.

Bug Fixes

Unicode v1.21.1

This is the changelog for Unicode v1.21.1 released on March 16th, 2026. For older changelogs please consult the release tag on GitHub

Bug Fixes

  • Compiles without warning on Elixir 1.20.0-rc.3.

Unicode v1.21.0

This is the changelog for Unicode v1.21.0 released on January 19th, 2025. For older changelogs please consult the release tag on GitHub

Enhancements

Unicode v1.20.0

This is the changelog for Unicode v1.20.0 released on September 11, 2024. For older changelogs please consult the release tag on GitHub

Enhancements

Unicode v1.19.0

This is the changelog for Unicode v1.19.0 released on February 29th, 2024. For older changelogs please consult the release tag on GitHub

Bug Fixes

Enhancements

  • Add indic_conjunc_break as a property, it is new in Unicode 15.1.

  • Fix performance regression in Unicode.replace_invalid/2 from the original UniRecover library and confirm the code memory usage remains a constant 128 bytes for all benchmark scenarios. Thanks to @Moosieus for the fabulous PR. Closes #10.

  • Unicode.replace_invalid(string, :utf8, replacement) delegates to String.replace_invalid/2 where available (which will be from Elixir 1.16 onwards).

  • Confirm that the README installation version matches the code version. Thanks to @Moosieus for the PR.

Unicode v1.18.0

This is the changelog for Unicode v1.18.0 released on October 19th, 2023. For older changelogs please consult the release tag on GitHub

Enhancements

  • Adds Unicode.replace_invalid/3 to force-validate a binary as a UTF string. Any of the UTF encodings may be validated. Any invalid codepoints or incomplete sequences are replaced with a replacement string. Many thanks to @Moosieus for the contribution.

Unicode v1.17.0

This is the changelog for Unicode v1.17.0 released on September 17th, 2023. For older changelogs please consult the release tag on GitHub

Enhancements

  • Updates to Unicode 15.1 data.

  • Improve the security of the mix unicode.download task.

Unicode v1.16.2

This is the changelog for Unicode v1.16.2 released on August 16th, 2023. For older changelogs please consult the release tag on GitHub

Enhancements

  • Change the parsing of "SpecialCasing.text" specifically to support casing in unicode_string/.

Unicode v1.16.1

This is the changelog for Unicode v1.16.1 released on April 22nd, 2023. For older changelogs please consult the release tag on GitHub

Big Fixes

Unicode v1.16.0

This is the changelog for Unicode v1.16.0 released on March 18th, 2023. For older changelogs please consult the release tag on GitHub

Enhancements

Unicode v1.15.0

This is the changelog for Unicode v1.15.0 released on September 17th, 2022. For older changelogs please consult the release tag on GitHub

Note there is no release 1.14. The release is 1.15 to align with Unicode 15 and it is expected to keep this pattern into the future.

Deprecations

  • Fix deprecation warnings for Elixir 1.14. Now requires Elixir 1.11 as a minimum release.

Enhancements

Bug Fixes

  • Fix code fences in docs to be Elixir

Unicode v1.13.1

This is the changelog for Unicode v1.13.1 released on September 16th, 2021. For older changelogs please consult the release tag on GitHub

Bug Fixes

  • When looking up scripts, general categories and properties we indirect through an alias table. But not all entries have alias so in the case alias lookup fails we still need to lookup using the original key.

Unicode v1.13.0

This is the changelog for Unicode v1.13.0 released on September 15th, 2021. For older changelogs please consult the release tag on GitHub

Enhancements

  • Change the application name to :unicode (in collaboration with @Qqwy). The old name ex_unicode will be retired.

Unicode v1.12.0

This is the changelog for Unicode v1.12.0 released on September 14th, 2021. For older changelogs please consult the release tag on GitHub

Enhancements

Unicode v1.12.0-rc.0

This is the changelog for Unicode v1.12.0-rc.0 released on August 27th, 2021. For older changelogs please consult the release tag on GitHub

Enhancements

  • Updates to Unicode 14 preview data.

Unicode v1.11.2

This is the changelog for Unicode v1.11.2 released on May 25th, 2021. For older changelogs please consult the release tag on GitHub

Bug fixes

  • Make ex_doc and benchee optional. Thanks to @fireproofsocks.

Unicode v1.11.1

This is the changelog for Unicode v1.11.1 released on January 5th, 2021. For older changelogs please consult the release tag on GitHub

Bug fixes

  • Restrict ex_doc to only :dev and :release. Closes #3. Thanks to @manuelmontenegro.

  • Fix spec for Unicode.all/0 so dialyzer is happy

Unicode v1.11.0

This is the changelog for Unicode v1.11.0 released on October 8th, 2020. For older changelogs please consult the release tag on GitHub

Bug fixes

  • Rename the derived category :visible to :graph and change the definition to that in Unicode Regular Expressions. Deprecate the derived category :visible.

Unicode v1.10.0

This is the changelog for Unicode v1.10.0 released on October 5th, 2020. For older changelogs please consult the release tag on GitHub

Bug fixes

  • Revert "Change the definition of the derived property All to be the disjoint set of unicode ranges, not the closed set." since All in the ICU means the full range of codepoints, assigned or otherwise.

  • Add :inets and :public_key to :extra_applicatons to avoid warnings on Elixir 1.11.

Enhancements

Unicode v1.9.0

This is the changelog for Unicode v1.9.0 released on October 4th, 2020. For older changelogs please consult the release tag on GitHub

Enhancements

  • Change the definition of the derived property All to be the disjoint set of unicode ranges, not the closed set.

Unicode v1.8.0

This is the changelog for Unicode v1.8.0 released on July 12th, 2020. For older changelogs please consult the release tag on GitHub

Enhancements

  • Add the east asian width property to Unicode.Property.fetch/2 API

  • Add the word break property to Unicode.Property.fetch/2 API

Unicode v1.7.0

This is the changelog for Unicode v1.7.0 released on June 22nd, 2020. For older changelogs please consult the release tag on GitHub

Enhancements

  • Add the emoji properties to Uniccode.Property.fetch/2 API

  • Add certificate verification to download process

Unicode v1.6.0

This is the changelog for Unicode v1.6.0 released on May 17th, 2020. For older changelogs please consult the release tag on GitHub

Enhancements

  • Add Unicode.Utils.case_folding/0

  • Add Unicode.Utils.special_casing/0

Unicode v1.5.0

This is the changelog for Unicode v1.5.0 released on March 14th, 2020. For older changelogs please consult the release tag on GitHub

Enhancements

  • Add derived categories :printable: and :visible:. :printable: implements the same semnantics String.printable?/1. :visible: combines the categories [[:L:][:N:][:M:][:P:][:S:][:Zs:]].

Unicode v1.4.1

This is the changelog for Unicode v1.4.1 released on March 11th, 2020. For older changelogs please consult the release tag on GitHub

Bug Fixes

  • Regenerate the assigned ranges for Unicode 13

Unicode v1.4.0

This is the changelog for Unicode v1.4.0 released on March 11th, 2020. For older changelogs please consult the release tag on GitHub

Enhancements

Updates Unicode to version 13.0.

As of March 2020, Unicode has introduced Unicode 13.0 and this data now forms the basis of ex_unicode version 1.40. Version 13 of Unicode adds 5,390 characters, for a total of 143,859 characters. These additions include four new scripts, for a total of 154 scripts, as well as 55 new emoji characters.

Adds derived categories for various quotation marks.

Although the unicode character database has a flag to indicate if a given codepoint is a quotation mark, the list does not include CJK quotation marks, dingbats or alternative encodings. Some additional derived categories are therefore added that are taken from Wikipedia. The added dervived categories are:

  • QuoteMark - all quote marks
  • QuoteMarkLeft - all quote marks used on the left
  • QuoteMarkRight - quote marks used on the right
  • QuoteMarkAmbidextrous - quote marks used either left or right
  • QuoteMarkSingle - single quote marks
  • QuoteMarkDouble - double quote marks

These additional derived categories can be used in Unicode Sets, for example:

iex> Unicode.Set.match? ?', "[[:quote_mark:]]"
true
iex> Unicode.Set.match? ?', "[[:quote_mark_left:]]"
false
iex> Unicode.Set.match? ?', "[[:quote_mark_ambidextrous:]]"
true

Unicode v1.3.1

This is the changelog for Unicode v1.3.1 released on January 8th, 2020. For older changelogs please consult the release tag on GitHub

Bug Fixes

  • Remove call to Code.ensure_compiled?/1 which is deprecated in Elixir 1.10.0.

  • Fix the ranges for the General Category :assigned.

Unicode v1.3.0

This is the changelog for Unicode v1.3.0 released on December 3rd, 2019. For older changelogs please consult the release tag on GitHub

Breaking Changes

  • Changed two module names: Unicode.Category becomes Unicode.GeneralCategory and Unicode.CombiningClass becomes Unicode.CanonicalCombiningClass. These names map directly to the Unicode standard names. It also means all property module names can be derived from the Unicode property name which is what the new Unicode.servers/0 function does.

Enhancements

  • Add property modules for line break, sentence break, grapheme cluster break and indic syllabic category. These properties are used by the CLDR and Unicode segmentation rules.

  • Add Unicode.servers/0 that maps property names and aliases to a module name that serves that property.

Bug fixes

  • Fixes Unicode.aliases/0 to correctly use the aliases in data/property_alias.txt

Unicode v1.2.0

This is the changelog for Unicode v1.2.0 released on November 27th, 2019. For older changelogs please consult the release tag on GitHub

Breaking Changes

  • Script names are now atoms instead of strings to be consistent with other properties

Enhancements

  • Add aliases/0, fetch/1 and get/1 to Unicode.Property

  • Added additional properties to Unicode.Property. The set now includes those from the UCD files DerivedCoreProperties.txt and PropList.txt.

Unicode v1.1.0

This is the changelog for Unicode v1.1.0 released on November 23rd, 2019. For older changelogs please consult the release tag on GitHub

Breaking Changes

Enhancements

  • Unicode.Category.categories/0 now returns the super categories as well as the subcategories. These super categories are computed at compile time by consolidating the relevant subcategories. Unicode.Category.category/1 will only return one category, and it will be the subcategory as it consistent with earlier releases.

  • Add Unicode.ranges/0 that returns all Unicode codepoints as a list of 2-tuples representing the disjoint ranges of valid codepoints. The list is in sorted order.

  • Add aliases/0 for Unicode.Category, Unicode.Script, Unicode.Block, and Unicode.CombiningClass which returns the alias map for the relevant module.

  • Add fetch/1 and get/1 for Unicode.Category, Unicode.Script, Unicode.Block, and Unicode.CombiningClass. These functions leverage Unicode property value aliases for retrieving codepoints.

  • Add Unicode.fetch_property/1 and Unicode.get_property/1 that return the module responsible for handling a given Unicode property.

  • Add Unicode.compact_ranges/1 that given a list of 2-tuple ranges will compact them into as small a list of contiguous blocks as possible

  • Documented all public functions

Unicode v1.0.0

This is the changelog for Unicode v1.0.0 released on November 14th, 2019. For older changelogs please consult the release tag on GitHub

Breaking Changes

  • Rename the module prefix to Unicode since this package is not linked in any way to the Cldr family. The hex package is renamed to ex_unicode.

Cldr Unicode v0.7.1

This is the changelog for Unicode v0.7.1 released on November 12th, 2019. For older changelogs please consult the release tag on GitHub

Bug Fixes

  • Fixes count/1 for blocks, scripts and categories

  • Replace deprecated String.normalize/2 with :unicode.characters_to_nfd_binary/ for OTP release 20 and later.

Cldr Unicode v0.7.0

This is the changelog for Unicode v0.7.0 released on November 12th, 2019. For older changelogs please consult the release tag on GitHub

Enhancements

  • Add is_whitespace/1 guard generator

Cldr Unicode v0.6.0

This is the changelog for Unicode v0.6.0 released on October 22nd, 2019. For older changelogs please consult the release tag on GitHub

Enhancements

  • Update to Emoji 12.1

Cldr Unicode v0.5.0

This is the changelog for Unicode v0.5.0 released on May 12th, 2019. For older changelogs please consult the release tag on GitHub

Enhancements

  • Update to Unicode 12.1

Cldr Unicode v0.4.0

This is the changelog for Unicode v0.4.0 released on April 30th, 2019. For older changelogs please consult the release tag on GitHub

Enhancements

  • Adds Cldr.Unicode.unaccent/1

Breaking Changes

  • Block names are now atoms instead of strings

Cldr Unicode v0.3.0

This is the changelog for Unicode v0.3.0 released on March 28th, 2019. For older changelogs please consult the release tag on GitHub

Enhancements

  • Updated to Unicode version 12

Cldr Unicode v0.2.0

This is the changelog for Unicode v0.2.0 released on February 24th, 2019. For older changelogs please consult the release tag on GitHub

Enhancements

  • Moves the public API to the Cldr.Unicode module.

  • Updates and adds documentation to all public functions.

  • Removes the text annotations from the compiled functions which materially reduces the size of the beam files.

Cldr Unicode v0.1.0

This is the changelog for Unicode v0.1.0 released on February 23rd, 2019. For older changelogs please consult the release tag on GitHub

Enhancements

  • Initial release