Unicode v2.2.0
This is the changelog for Unicode v2.2.0 released on September 20th, 2026. For older changelogs please consult the release tag on GitHub
This release makes every property value the Unicode Character Database declares resolvable, including the values a data file supplies only through an @missing annotation.
Enhancements
Resolves the default value of fifteen further properties, so
\p{jt=U},\p{bpt=None},\p{sc=Unknown},\p{age=NA}and their companions now answer. Every property's values together account for all 1,114,112 codepoints, where before only five did.Applies the range-specific defaults
BidiClass.txtdeclares, so an unassigned codepoint in a right-to-left block resolves asRorALrather thanL.Resolves
Joining_Groupvalues from any spelling, including theCrown_AinformPropertyValueAliases.txtuses. 73 of the 116 values previously resolved from neither spelling.Resolves binary property aliases that contain a separator, such as
Bidi_M,Gr_BaseandPat_Syn, which the lookup normalised but the alias table did not.Documents
Unicode.fetch_property/1andUnicode.get_property/1, which resolve a property named at runtime to the module that serves it. Both were public but carried@doc false, so the lookup thatunicode_setevaluates set expressions with was absent from the documentation.Adds an introduction guide covering the shape of the API: the three ways to reach a property, naming property values, character names, script extensions and the guards. Its examples run as doctests.
Loads
NameAliases.txt, soUnicode.CharacterName.to_codepoint/2resolves the names of the control characters, which have noNameproperty, along with abbreviations such asLFand the corrections for characters whose published name contains an error.Unicode.CharacterName.aliases/1returns a codepoint's aliases with their types.
Bug fixes
Unicode.EastAsianWidth.east_asian_width_category/1returns the UCD default:n, rather than:otherwhich is not a value of the property, for codepointsEastAsianWidth.txtomits.
Notes
known_*/0andcount/1include the default value for the affected properties, so those lists and counts are larger than in 2.1.0.
Unicode v2.1.0
This release closes a number of Unicode Character Database coverage gaps: several enumerated properties that appeared in PropertyAliases.txt but had no backing data are now resolvable through Unicode.fetch_property/1 and have their own introspection modules.
Enhancements
Resolves the property value that a UCD
@missingannotation declares as the default.Grapheme_Cluster_Break,Word_BreakandSentence_Breaknow answerOther(and itsXXalias), andLine_BreakanswersXX, for every codepoint the data files do not list.Line_Breakpreviously covered only 358,578 of the 1,114,112 codepoints; the remainder, including all private use, resolved only through a lookup fallback and was invisible to anything enumerating the property. Where a value is both listed explicitly and named as the@missingdefault, asXXis forLine_Break, the two sets are merged rather than replaced.Adds the
ageproperty value aliases fromPropertyValueAliases.txt, soUnicode.Age.fetch("V18_0")resolves the same set asUnicode.Age.fetch("18.0").Updates the underlying data to the final Unicode 18.0 release (16 September 2026), adding the
Jurchen,Proto_CuneiformandSealscripts, seven new blocks and ten newCrown_*joining groups.Adds
Unicode.ScriptExtensionsfor theScript_Extensions(scx) property, the last UCD enumerated property without a backing module and a UTS #18 RL1.2 conformance requirement.Unicode.ScriptExtensions.script_extensions/1returns the set of scripts a codepoint is used with, defaulting to itsScriptvalue where the UCD lists no explicit set.Unicode.CharacterName.to_codepoint/2accepts a:fuzzyoption that resolves a misspelled name byString.jaro_distance/2, either astruefor the default distance of0.8or as an explicit distance. It succeeds only where one name is strictly closest than all others, so an ambiguous query returns:errorrather than guessing.Adds
Unicode.CharacterName.to_name/1, the reverse ofto_codepoint/1, returning theNameproperty of a codepoint. Each name is still stored only once — the original is reconstructed from the normalized form plus separator positions — so reverse lookup adds about 410KB rather than a second name table, and leavesto_codepoint/1unchanged in speed.Unicode.CharacterName.to_codepoint/1now resolves algorithmically derived names — CJK and Tangut ideographs, Hangul syllables, and the Seal and Jurchen characters — covering a further 131,576 characters. The names are computed from the UAX #44 derivation rules on a table-lookup miss rather than stored, adding under 4KB.mix unicode.downloadnow resolves the UCD and emoji trees through independent release channels, settable per run with--release/--emoji-release, by environment variable, or in application config. Adds--intoto download to a scratch directory,--dry-run, and a post-download check that every data file reports the same Unicode version.Adds the enumerated property modules
Unicode.Age,Unicode.NumericType,Unicode.NumericValue,Unicode.DecompositionType,Unicode.HangulSyllableType,Unicode.IndicPositionalCategory,Unicode.VerticalOrientation,Unicode.JoiningGroupandUnicode.BidiPairedBracketType, each with a codepoint lookup, range introspection and a top-level delegate onUnicode.Adds the normalization quick check modules
Unicode.NfcQuickCheck,Unicode.NfdQuickCheck,Unicode.NfkcQuickCheckandUnicode.NfkdQuickCheck(theNFC_QC,NFD_QC,NFKC_QCandNFKD_QCproperties) withUnicode.nfc_quick_check/1and friends.Adds
Bidi_Mirroredas a first class boolean property, soUnicode.Property.bidi_mirrored?/1andUnicode.properties/1now report it.The
Nameproperty now resolves throughUnicode.fetch_property/1toUnicode.CharacterName.Adds the UTS #58 link properties
Unicode.LinkTerm,Unicode.LinkEmailandUnicode.LinkBracket, which together drive link detection in flowing text.mix unicode.downloadgains alinkificationtree alongside the UCD and emoji trees to fetch them.Adds the
LC(Cased_Letter) general category, the union ofLu,LlandLt. It is the only General_Category group whose name is not a single letter, so it was not produced by the derivation that yieldsL,N,Pand the rest, andUnicode.GeneralCategory.fetch("lc")returned:errordespite the alias being advertised.
Bug fixes
Property aliases containing an underscore (for example
nfc_qc) are now reachable throughUnicode.fetch_property/1, which previously only matched the whitespace-stripped canonical form.Scripts and blocks whose
PropertyValueAliasesline carries more than two names now resolve by every spelling.Unicode.Script.fetch/1failed for"copt"and"qaac", andUnicode.Block.fetch/1for"latin1sup"and"cyrillicsupplementary", because the alias map was built by inverting a many-to-one map and elected a name absent from the data.Unicode.version/0no longer reports a stale version after the data files are updated. It derives the version fromblocks.txtat compile time but did not declare it as an@external_resource, so the module was not recompiled when the data changed.mix unicode.downloadno longer requests a non-existent emoji URL. The emoji files are versioned separately from the UCD andPublic/emoji/17.0/was never published, so the previous version-interpolated path returned 404 and the files had to be updated by hand.
Unicode v2.0.0
This is the changelog for Unicode v2.0.0 released on July 9th, 2026. For older changelogs please consult the release tag on GitHub
This is a major release that folds the unicode_guards library into unicode and replaces the hand-maintained derived category tables with values computed directly from the character database.
Breaking changes
The
Unicode.Guardsmodule is now provided by this library. The separateunicode_guardspackage is no longer required; remove it from your dependencies and depend onunicodeinstead.The derived general categories
:Assigned,:Graph,:Visibleand:Printablenow reflect the current Unicode data. The previous static tables were stale::Assignedgained 15,942 codepoints (283,440 to 299,382),:Graph/:Visiblegained the same 15,942 (281,308 to 297,250), and:Printableis corrected to matchString.printable?/1(it previously excluded most of the BMP, including Arabic, CJK and Hangul).The internal
Unicode.DerivedCategory.AssignedandUnicode.DerivedCategory.Graphmodules have been removed; these categories are now computed at compile time.
Enhancements
Adds
Unicode.CharacterName.to_codepoint/1, which resolves a Unicode character name to its codepoint with loose matching. The name table is stored as a sorted binary blob with a binary search to keep it compact.Codepoint lookup functions such as
Unicode.GeneralCategory.category/1,Unicode.Script.script/1and theUnicode.Propertyboolean functions are now implemented with binary search over compact range tables instead of very large generated guard clauses. Compilation is an order of magnitude faster and codepoint lookups are approximately 10x faster.Adds
Unicode.RangeSearchwhich builds the range search tables at compile time and performs the binary search over them.Documentation for all public modules and functions now follows a standard format with arguments, return values and examples.
Removes the unused and incorrect
Unicode.Utils.remove_reserved_codepoints/1.Adds
Unicode.Guards, a set of guards (is_upper/1,is_lower/1,is_digit/1,is_whitespace/1,is_graph/1, the quotation-mark guards and more) for use in functionwhenclauses. Folded in fromunicode_guardswith no runtime dependencies.Derived categories are computed from the character database using new range-set helpers (
Unicode.Utils.union_ranges/1,complement_ranges/1anddifference_ranges/2), removing the need to regenerate static tables by hand each release.
Bug Fixes
Fix UTF-16 and UTF-32 validation.
Unicode.replace_invalid/3for the:utf16,:utf16be,:utf16le,:utf32,:utf32beand:utf32leencodings previously crashed on any input and dropped valid codepoints.Fix
:utf32validation dispatching to the UTF-16 implementation.Replacement strings for UTF-16 and UTF-32 validation are now transcoded to the target encoding rather than being spliced in as UTF-8 bytes.
Fix
Unicode.compact_ranges/1truncating a range that fully contains the following range.Fix
Unicode.Emoji.emoji/1returningnilfor single-codepoint emoji graphemes in a string.Fix
Unicode.Block.fetch/1andUnicode.Block.get/1for block names containing digits, such as"Latin-1 Supplement"and"Number Forms", which previously returned:error/nilfor every spelling of the canonical name because no alias mapped the normalised name to the block key.
Unicode v1.22.0
This is the changelog for Unicode v1.22.0 released on May 4th, 2026. For older changelogs please consult the release tag on GitHub
Enhancements
Adds the
Bidi_ClassUnicode property.Adds the
Joining_TypeUnicode property.
Unicode v1.21.2
This is the changelog for Unicode v1.21.2 released on April 29th, 2026. For older changelogs please consult the release tag on GitHub
Breaking changes
- Supported on Elixir 1.17 and later only.
Bug Fixes
- Fix type on
Unicode.script/1
Unicode v1.21.1
This is the changelog for Unicode v1.21.1 released on March 16th, 2026. For older changelogs please consult the release tag on GitHub
Bug Fixes
- Compiles without warning on Elixir 1.20.0-rc.3.
Unicode v1.21.0
This is the changelog for Unicode v1.21.0 released on January 19th, 2025. For older changelogs please consult the release tag on GitHub
Enhancements
- Updates to Unicode 17.0 data.
Unicode v1.20.0
This is the changelog for Unicode v1.20.0 released on September 11, 2024. For older changelogs please consult the release tag on GitHub
Enhancements
- Updates to Unicode 16.0 data.
Unicode v1.19.0
This is the changelog for Unicode v1.19.0 released on February 29th, 2024. For older changelogs please consult the release tag on GitHub
Bug Fixes
Unicode.properties/1no longer does anEnum.uniq/1on the result since that breaks the functions contract.Fix
Unicode.Emoji.emoji/0.Fix documentation in
Unicode.WordBreak.
Enhancements
Add
indic_conjunc_breakas a property, it is new in Unicode 15.1.Fix performance regression in
Unicode.replace_invalid/2from the original UniRecover library and confirm the code memory usage remains a constant 128 bytes for all benchmark scenarios. Thanks to @Moosieus for the fabulous PR. Closes #10.Unicode.replace_invalid(string, :utf8, replacement)delegates toString.replace_invalid/2where available (which will be from Elixir 1.16 onwards).Confirm that the README installation version matches the code version. Thanks to @Moosieus for the PR.
Unicode v1.18.0
This is the changelog for Unicode v1.18.0 released on October 19th, 2023. For older changelogs please consult the release tag on GitHub
Enhancements
- Adds
Unicode.replace_invalid/3to force-validate a binary as a UTF string. Any of the UTF encodings may be validated. Any invalid codepoints or incomplete sequences are replaced with a replacement string. Many thanks to @Moosieus for the contribution.
Unicode v1.17.0
This is the changelog for Unicode v1.17.0 released on September 17th, 2023. For older changelogs please consult the release tag on GitHub
Enhancements
Updates to Unicode 15.1 data.
Improve the security of the
mix unicode.downloadtask.
Unicode v1.16.2
This is the changelog for Unicode v1.16.2 released on August 16th, 2023. For older changelogs please consult the release tag on GitHub
Enhancements
- Change the parsing of "SpecialCasing.text" specifically to support casing in unicode_string/.
Unicode v1.16.1
This is the changelog for Unicode v1.16.1 released on April 22nd, 2023. For older changelogs please consult the release tag on GitHub
Big Fixes
- Fix spelling of
Unicode.script_dominance/1.
Unicode v1.16.0
This is the changelog for Unicode v1.16.0 released on March 18th, 2023. For older changelogs please consult the release tag on GitHub
Enhancements
Add
Unicode.script_statistic/1that returns the first index and grapheme count of the scripts in a string. This is useful to help derive the likely locale of a string. Determining the locale is outside the scope of this library but is required in ex_cldr_person_names.Add
Unicode.script_dominance/1to sort the results ofUnicode.script_statistic/1in descending dominance order.
Unicode v1.15.0
This is the changelog for Unicode v1.15.0 released on September 17th, 2022. For older changelogs please consult the release tag on GitHub
Note there is no release 1.14. The release is 1.15 to align with Unicode 15 and it is expected to keep this pattern into the future.
Deprecations
- Fix deprecation warnings for Elixir 1.14. Now requires Elixir 1.11 as a minimum release.
Enhancements
- Updates to Unicode 15.
Bug Fixes
- Fix code fences in docs to be Elixir
Unicode v1.13.1
This is the changelog for Unicode v1.13.1 released on September 16th, 2021. For older changelogs please consult the release tag on GitHub
Bug Fixes
- When looking up scripts, general categories and properties we indirect through an alias table. But not all entries have alias so in the case alias lookup fails we still need to lookup using the original key.
Unicode v1.13.0
This is the changelog for Unicode v1.13.0 released on September 15th, 2021. For older changelogs please consult the release tag on GitHub
Enhancements
- Change the application name to
:unicode(in collaboration with @Qqwy). The old nameex_unicodewill be retired.
Unicode v1.12.0
This is the changelog for Unicode v1.12.0 released on September 14th, 2021. For older changelogs please consult the release tag on GitHub
Enhancements
- Update to use Unicode 14 release data.
Unicode v1.12.0-rc.0
This is the changelog for Unicode v1.12.0-rc.0 released on August 27th, 2021. For older changelogs please consult the release tag on GitHub
Enhancements
- Updates to Unicode 14 preview data.
Unicode v1.11.2
This is the changelog for Unicode v1.11.2 released on May 25th, 2021. For older changelogs please consult the release tag on GitHub
Bug fixes
- Make
ex_docandbencheeoptional. Thanks to @fireproofsocks.
Unicode v1.11.1
This is the changelog for Unicode v1.11.1 released on January 5th, 2021. For older changelogs please consult the release tag on GitHub
Bug fixes
Restrict
ex_docto only:devand:release. Closes #3. Thanks to @manuelmontenegro.Fix spec for
Unicode.all/0so dialyzer is happy
Unicode v1.11.0
This is the changelog for Unicode v1.11.0 released on October 8th, 2020. For older changelogs please consult the release tag on GitHub
Bug fixes
- Rename the derived category
:visibleto:graphand change the definition to that in Unicode Regular Expressions. Deprecate the derived category:visible.
Unicode v1.10.0
This is the changelog for Unicode v1.10.0 released on October 5th, 2020. For older changelogs please consult the release tag on GitHub
Bug fixes
Revert "Change the definition of the derived property
Allto be the disjoint set of unicode ranges, not the closed set." sinceAllin the ICU means the full range of codepoints, assigned or otherwise.Add
:inetsand:public_keyto:extra_applicatonsto avoid warnings on Elixir 1.11.
Enhancements
Add
Unicode.assigned/0to return the list of codepoint ranges that are assigned within UnicodeRename
Unicode.ranges/0toUnicode.all/0to better reflect the intent.Unicode.ranges/0is deprecated.
Unicode v1.9.0
This is the changelog for Unicode v1.9.0 released on October 4th, 2020. For older changelogs please consult the release tag on GitHub
Enhancements
- Change the definition of the derived property
Allto be the disjoint set of unicode ranges, not the closed set.
Unicode v1.8.0
This is the changelog for Unicode v1.8.0 released on July 12th, 2020. For older changelogs please consult the release tag on GitHub
Enhancements
Add the east asian width property to
Unicode.Property.fetch/2APIAdd the word break property to
Unicode.Property.fetch/2API
Unicode v1.7.0
This is the changelog for Unicode v1.7.0 released on June 22nd, 2020. For older changelogs please consult the release tag on GitHub
Enhancements
Add the emoji properties to
Uniccode.Property.fetch/2APIAdd certificate verification to download process
Unicode v1.6.0
This is the changelog for Unicode v1.6.0 released on May 17th, 2020. For older changelogs please consult the release tag on GitHub
Enhancements
Add
Unicode.Utils.case_folding/0Add
Unicode.Utils.special_casing/0
Unicode v1.5.0
This is the changelog for Unicode v1.5.0 released on March 14th, 2020. For older changelogs please consult the release tag on GitHub
Enhancements
- Add derived categories
:printable:and:visible:.:printable:implements the same semnanticsString.printable?/1.:visible:combines the categories[[:L:][:N:][:M:][:P:][:S:][:Zs:]].
Unicode v1.4.1
This is the changelog for Unicode v1.4.1 released on March 11th, 2020. For older changelogs please consult the release tag on GitHub
Bug Fixes
- Regenerate the assigned ranges for Unicode 13
Unicode v1.4.0
This is the changelog for Unicode v1.4.0 released on March 11th, 2020. For older changelogs please consult the release tag on GitHub
Enhancements
Updates Unicode to version 13.0.
As of March 2020, Unicode has introduced Unicode 13.0 and this data now forms the basis of ex_unicode version 1.40. Version 13 of Unicode adds 5,390 characters, for a total of 143,859 characters. These additions include four new scripts, for a total of 154 scripts, as well as 55 new emoji characters.
Adds derived categories for various quotation marks.
Although the unicode character database has a flag to indicate if a given codepoint is a quotation mark, the list does not include CJK quotation marks, dingbats or alternative encodings. Some additional derived categories are therefore added that are taken from Wikipedia. The added dervived categories are:
- QuoteMark - all quote marks
- QuoteMarkLeft - all quote marks used on the left
- QuoteMarkRight - quote marks used on the right
- QuoteMarkAmbidextrous - quote marks used either left or right
- QuoteMarkSingle - single quote marks
- QuoteMarkDouble - double quote marks
These additional derived categories can be used in Unicode Sets, for example:
iex> Unicode.Set.match? ?', "[[:quote_mark:]]"
true
iex> Unicode.Set.match? ?', "[[:quote_mark_left:]]"
false
iex> Unicode.Set.match? ?', "[[:quote_mark_ambidextrous:]]"
trueUnicode v1.3.1
This is the changelog for Unicode v1.3.1 released on January 8th, 2020. For older changelogs please consult the release tag on GitHub
Bug Fixes
Remove call to
Code.ensure_compiled?/1which is deprecated in Elixir 1.10.0.Fix the ranges for the General Category
:assigned.
Unicode v1.3.0
This is the changelog for Unicode v1.3.0 released on December 3rd, 2019. For older changelogs please consult the release tag on GitHub
Breaking Changes
- Changed two module names:
Unicode.CategorybecomesUnicode.GeneralCategoryandUnicode.CombiningClassbecomesUnicode.CanonicalCombiningClass. These names map directly to the Unicode standard names. It also means all property module names can be derived from the Unicode property name which is what the newUnicode.servers/0function does.
Enhancements
Add property modules for line break, sentence break, grapheme cluster break and indic syllabic category. These properties are used by the CLDR and Unicode segmentation rules.
Add
Unicode.servers/0that maps property names and aliases to a module name that serves that property.
Bug fixes
- Fixes
Unicode.aliases/0to correctly use the aliases indata/property_alias.txt
Unicode v1.2.0
This is the changelog for Unicode v1.2.0 released on November 27th, 2019. For older changelogs please consult the release tag on GitHub
Breaking Changes
- Script names are now atoms instead of strings to be consistent with other properties
Enhancements
Add
aliases/0,fetch/1andget/1toUnicode.PropertyAdded additional properties to
Unicode.Property. The set now includes those from the UCD filesDerivedCoreProperties.txtandPropList.txt.
Unicode v1.1.0
This is the changelog for Unicode v1.1.0 released on November 23rd, 2019. For older changelogs please consult the release tag on GitHub
Breaking Changes
- Removed
Unicode.Guardsfrom this library and moved them to the unicode_set package.
Enhancements
Unicode.Category.categories/0now returns the super categories as well as the subcategories. These super categories are computed at compile time by consolidating the relevant subcategories.Unicode.Category.category/1will only return one category, and it will be the subcategory as it consistent with earlier releases.Add
Unicode.ranges/0that returns all Unicode codepoints as a list of 2-tuples representing the disjoint ranges of valid codepoints. The list is in sorted order.Add
aliases/0forUnicode.Category,Unicode.Script,Unicode.Block, andUnicode.CombiningClasswhich returns the alias map for the relevant module.Add
fetch/1andget/1forUnicode.Category,Unicode.Script,Unicode.Block, andUnicode.CombiningClass. These functions leverage Unicode property value aliases for retrieving codepoints.Add
Unicode.fetch_property/1andUnicode.get_property/1that return the module responsible for handling a given Unicode property.Add
Unicode.compact_ranges/1that given a list of 2-tuple ranges will compact them into as small a list of contiguous blocks as possibleDocumented all public functions
Unicode v1.0.0
This is the changelog for Unicode v1.0.0 released on November 14th, 2019. For older changelogs please consult the release tag on GitHub
Breaking Changes
- Rename the module prefix to
Unicodesince this package is not linked in any way to theCldrfamily. The hex package is renamed toex_unicode.
Cldr Unicode v0.7.1
This is the changelog for Unicode v0.7.1 released on November 12th, 2019. For older changelogs please consult the release tag on GitHub
Bug Fixes
Fixes
count/1for blocks, scripts and categoriesReplace deprecated
String.normalize/2with:unicode.characters_to_nfd_binary/for OTP release 20 and later.
Cldr Unicode v0.7.0
This is the changelog for Unicode v0.7.0 released on November 12th, 2019. For older changelogs please consult the release tag on GitHub
Enhancements
- Add
is_whitespace/1guard generator
Cldr Unicode v0.6.0
This is the changelog for Unicode v0.6.0 released on October 22nd, 2019. For older changelogs please consult the release tag on GitHub
Enhancements
- Update to Emoji 12.1
Cldr Unicode v0.5.0
This is the changelog for Unicode v0.5.0 released on May 12th, 2019. For older changelogs please consult the release tag on GitHub
Enhancements
- Update to Unicode 12.1
Cldr Unicode v0.4.0
This is the changelog for Unicode v0.4.0 released on April 30th, 2019. For older changelogs please consult the release tag on GitHub
Enhancements
- Adds
Cldr.Unicode.unaccent/1
Breaking Changes
- Block names are now atoms instead of strings
Cldr Unicode v0.3.0
This is the changelog for Unicode v0.3.0 released on March 28th, 2019. For older changelogs please consult the release tag on GitHub
Enhancements
- Updated to Unicode version 12
Cldr Unicode v0.2.0
This is the changelog for Unicode v0.2.0 released on February 24th, 2019. For older changelogs please consult the release tag on GitHub
Enhancements
Moves the public API to the
Cldr.Unicodemodule.Updates and adds documentation to all public functions.
Removes the text annotations from the compiled functions which materially reduces the size of the beam files.
Cldr Unicode v0.1.0
This is the changelog for Unicode v0.1.0 released on February 23rd, 2019. For older changelogs please consult the release tag on GitHub
Enhancements
- Initial release