When we receive an attribute value, we put it into a standard form before we use it to link events and build features. This means ABC123, abc123 and ␣abc123␣ count as the same customer token, so a customer's history isn't split across the small variations different apps and channels send.
Base standard
We use Normalization Form KD (NFKD), "Compatibility Decomposition", from Unicode Standard Annex #15: Unicode Normalization Forms. The standard describes the goal:
"When implementations keep strings in a normalized form, they can be assured that equivalent strings have a unique binary representation."
NFKD is the broadest of the four Unicode normalization forms. It splits accented letters into a base letter followed by its accent, and it folds look-alike characters into their plain equivalents. For example, full-width A becomes A and the circled digit ① becomes 1.
Our process adds steps before and after NFKD, so the result is not pure NFKD. The differences are listed under Differences from the standard below.
What happens to a text value in general
Text attributes such as the customer token, username and name fields go through these steps, in order:
- Remove every character that isn't a letter, a number, ASCII punctuation (the punctuation on a US keyboard) or a plain space. This drops tabs, line breaks, emoji, symbols such as
€and any punctuation outside ASCII. - Encode HTML characters so the value is safe to display:
'"<>&and--. - Apply NFKD.
- Convert to lowercase.
- Tidy spaces. Runs of spaces become one space, and leading and trailing spaces are removed.
Variations by attribute
| Attribute | Treatment |
|---|---|
| Customer token, username, first/middle/last name etc. | The five text steps above |
| Email address | As text, except spaces are removed and HTML characters aren't encoded. Then every dot before the @ is ignored, for every email provider. Anything after a + is kept |
| Phone number (E.164) | Reformatted to E.164 when the number starts with + and a country code. A number without a country code is used as sent |
| Country code | Converted to the two-letter ISO 3166 code |
| Currency code | Converted to the three-letter ISO 4217 code in uppercase |
| Bank account number | Leading zeros removed |
| IBAN | Spaces removed, uppercase |
| Date (for example date of birth) | Written as YYYY-MM-DD |
Custom fields (custom.general_purpose) |
Sent through the API, stored exactly as sent. If extracted or copied to by a journey, steps 3 and 5 only (case is kept) |
Examples
| Input | Becomes | What happened |
|---|---|---|
NormProbe-ABC-123 |
normprobe-abc-123 |
Lowercase |
␣␣Your␣␣␣Name␣␣ |
your name |
Spaces tidied |
abc⇥def↵ghi (tab, line break) |
abcdefghi |
Removed, not turned into spaces |
user😀01 |
user01 |
Emoji removed |
Test'User (straight apostrophe) |
test'user |
Apostrophe HTML-encoded (step 2) |
Test’User (curly apostrophe, the default on phone keyboards) |
testuser |
Removed as non-ASCII punctuation (step 1) |
José Müller |
josé müller |
Lowercase, accents kept |
ABC123 (full-width) |
abc123 |
Folded to standard characters |
①②③ |
123 |
Folded to standard digits |
ガギグ (half-width katakana) |
ガギグ |
Folded to full-width |
テスト ユーザー。 |
テストユーザー |
Ideographic space and 。 removed |
ทดสอบ ผู้ใช้ |
ทดสอบ ผูใช |
Thai tone mark removed |
ข้าว |
ขาว |
Thai tone mark ้ removed |
สวัสดีครับ |
สวัสดีครับ |
Thai vowel marks kept |
John.Smith+Tag@Example.COM (email) |
johnsmith+tag@example.com |
Lowercase, dots before @ ignored |
+61 (0)412-345-678 (phone) |
+61412345678 |
E.164 |
AUS or 036 (country) |
AU |
ISO 3166 two-letter code |
840 or aud (currency) |
USD, AUD |
ISO 4217 code |
0011223344 (account number) |
11223344 |
Leading zeros removed |
gb82 west 1234 5698 7654 32 (IBAN) |
GB82WEST12345698765432 |
Spaces removed, uppercase |
20270405 (date) |
2027-04-05 |
Standard date format |
Differences from the standard
For comparison we use Unicode's own definition of a caseless, compatibility-normalised match (The Unicode Standard, section 3.13, definition D146):
"A string X is a compatibility caseless match for a string Y if and only if: NFKD(toCasefold(NFKD(toCasefold(NFD(X))))) = NFKD(toCasefold(NFKD(toCasefold(NFD(Y)))))"
Our normalization differs this in three ways.
1. We remove characters before applying NFKD. The standard never removes characters. Because our removal step runs first, it also sees characters before NFKD has had a chance to rewrite them:
- Accents sent as a separate character are removed.
écan be sent as one character or asefollowed by a combining accent. The standard treats these as the same text. We keep the accent in the first case and remove it in the second, soJosécan become eitherjoséorjosedepending on how the app or keyboard encoded it. Japanese voiced marks behave the same way (ガorカ). - Thai tone marks are removed. Unicode doesn't classify tone marks (
่ ้ ๊ ๋), maitaikhu (็) or thanthakhat (์) as letters, so step 1 removes them. Vowel marks count as letters and are kept. No Unicode normalization form removes either. - Characters NFKD would have converted are removed instead. The standard turns these into ordinary characters that our process would keep. We remove them first.
| Input | Standard (D146) | Darwinium |
|---|---|---|
José with a separate accent |
josé |
jose |
ข้าว |
ข้าว |
ขาว |
ABC! (full-width !) |
abc! |
abc |
a…b |
a...b |
ab |
ab + non-breaking space + cd |
ab cd |
abcd |
東京 タワー (ideographic space) |
東京 タワー |
東京タワー |
㍿ |
株式会社 |
rejected (nothing left) |
㌔ |
キロ |
rejected (nothing left) |
2. We lowercase, where the standard case-folds. Case folding is Unicode's method for matching text regardless of case. Lowercasing gives the same result for most text, with these exceptions:
| Input | Standard (D146) | Darwinium |
|---|---|---|
Straße and STRASSE |
both strasse |
straße and strasse (different) |
ΟΔΟΣ and οδοσ |
both οδοσ |
οδος and οδοσ (different) |
3. We add steps the standard doesn't have. HTML encoding (step 2), space tidying (step 5), ignoring dots in email addresses, and the rejections below are Darwinium rules, not Unicode ones.
Some results are standard NFKD and can still be surprising:
- An accented letter is stored as the letter followed by the accent.
- Korean syllables are stored as their component letters (
한becomesᄒ+ᅡ+ᆫ). They still display as 한, so the change only shows up when you compare the bytes. - Thai
ำis split intoํ+า.
If you compare values exported from Darwinium with your own records, apply NFKD to your side first.
Values we reject
If a value is empty after step 1 (for example only spaces or only emoji), or looks like a payment card number, we drop that attribute. The event is still processed, the API responds with HTTP 422, and the response lists the attribute under invalid_input_attributes.
The card number check applies to values of 12 to 19 digits that start like a real card number and pass the Luhn check. If your customer tokens are plain numbers, add the attribute to confirmed_non_card_attributes in your journey to switch the check off for it.
Values sent through the API and values extracted by a journey are normalised the same way. Custom fields are the exception noted above.