Menu

Name Data by Locale: Why One Name Field Never Fits

Name data by locale is not a single field with two halves. Family-name order, single names, patronymics and diacritics all change what a form has to store.

Published

  • test data
  • identity
  • international

Name data by locale is where a form design that looked neutral turns out to have been written for one part of the world. The assumption is almost always the same — that a person has a given name followed by a family name, in that order, in the Latin alphabet, separated by a space — and every clause of that sentence is false somewhere.

This guide goes through the structural differences that matter when data is stored rather than displayed, and explains what a test suite has to contain before anyone can claim a name field is internationalised.

What a name is made of varies

The pieces that make up a personal name are not universal, and neither is their number. Common arrangements include a given name plus a family name, a family name written before a given name, one or more given names with no family name at all, a given name followed by a patronymic derived from a parent’s given name, and compound family names written with a space or a hyphen between the parts.

None of these is an exception to a rule; they are simply different rules. A form that reasons about “the first name” and “the last name” is asserting a structure that a substantial share of the world does not follow, and the assertion fails at the point where the data is used rather than where it is entered — in a greeting, in a shipping label, in a sorted list, in a matching algorithm.

Which languages put the family name first?

Several do, and they are not a small group. East Asian naming conventions conventionally place the family name first and the given name second, and Hungarian does the same within Europe. When such a name is recorded in a system that expects the opposite order, the two halves get swapped, and the swap is invisible because both halves are plausible personal names.

The failure becomes severe when the name is matched against another record or printed on a document. A letter addressed to the wrong half of a name is a small embarrassment; a border-control or identity document showing the wrong name order is a different category of problem altogether, and it is one of the reasons international standards for machine-readable documents exist.

The practical lesson is that the order is a property of the record, not a property of the field. Store the parts in a defined way, store the locale that governs their order, and format for display at the last moment rather than at the point of entry.

Are single names legitimate?

Yes, and a form that requires two name fields will reject a real person. Mononyms exist as legal names in several countries and cultures, and people who hold them run into this constantly: the form insists on a family name, so they type something, and now every document they are compared against disagrees with the record.

The same pattern appears in milder forms. Some people have several given names and no middle-name slot that fits them. Some have a family name composed of two words that a space-separated parser will split. Some have names that are a single character, which trips minimum-length checks written for reassurance rather than for a reason.

The fix is structural rather than cosmetic: mark the second name field as optional for the countries where it is optional, or better, treat the whole name as one value with optional part fields that are never required by default. Validation should reject what is impossible, not what is merely unfamiliar.

What happens to a name in another alphabet?

Two things happen, and they are often confused with each other. The first is transliteration: converting a name written in one writing system into another, which is a mapping with real choices in it and more than one accepted convention. The second is normalisation: deciding whether two strings that look identical are actually identical, which is a storage and comparison question.

Diacritics matter in both. A name carrying an accent is a different string from the same name without it, and whether they should be treated as equal depends on what you are doing. Matching two records should usually be accent-insensitive; printing a document should never strip accents. Collation — the order names sort in — is also language-specific, so a list sorted by the server’s default locale will not be in the order a reader of that language expects.

Case is subtler than it looks. Converting a name to upper case can change its length in some languages, and a small number of letter pairs have different upper-case forms depending on the language. A name field that upper-cases for display and stores the result has destroyed information it cannot reconstruct.

Why encoding issues keep resurfacing

Two strings can render identically on screen and still not be equal, because a character with an accent can be represented either as one precomposed character or as a base character followed by a combining mark. Both are valid, both display the same way, and comparing them byte for byte returns false.

The same problem appears whenever text moves between systems with different assumptions about what an identifier means. The durable fix is to normalise on the boundary — when data enters, when it is compared, and when it is exported — using one agreed form rather than whichever one arrived.

This is also why a test suite needs non-Latin and accented names in it. A fixture made entirely of plain Latin names will pass every check while the production data quietly breaks three of them. Records from the identity and test data generator are drawn per country and carry that country’s naming conventions, which makes them convenient for this kind of coverage; they are synthetic names for testing purposes only and do not belong to anyone.

How should name fields be laid out?

Start from what the data is used for and work backwards. Most systems need to display a name, sort a name, and search a name, and none of those operations requires the name to be split into exactly two parts at entry.

  • A single display field holds the name as it should be read, in the order the record’s locale dictates.
  • Separate part fields are optional and never both required; a form that insists on a family name is not international.
  • A locale or script marker on the record lets formatting and sorting make the right choice later.
  • Length limits are generous, because real names are longer than the examples in any specification.
  • Search indexes the normalised form, while the stored value keeps its original characters.

That layout costs one extra column and removes an entire class of defects. It also stops the name being silently reformatted by whichever layer happens to touch it first.

For developers: fields, order and comparison

Model the name as a small structure with a clear owner. Store the parts you were given, plus the order they belong in, and derive the display string on demand. Never store the result of a case conversion or an accent strip; store the original and normalise a copy for comparison.

For fixtures, cover the shapes that break assumptions rather than adding more ordinary ones: a mononym, a family name placed first, a hyphenated compound, a name with a combining-mark accent, and a name that is a single character long. Those five catch most defects that a hundred ordinary names will not.

Then check the two operations people forget. Sorting should follow the record’s locale rather than the machine’s, and truncation should be verified against the longest name the data actually contains, because a layout that fits every example in the specification is not evidence that it fits the database. Every name in a test fixture is invented for software testing, and no record generated this way identifies or stands in for a real person.

Next steps

Make the family-name field optional in one form this week and see whether anything downstream depends on it; if the answer is yes, the dependency is the bug. Then generate a set of names across several locales in the identity generator and run them through your display, sort and search paths, checking that sorting changes with the record’s locale instead of the server’s. The field consistency article covers the other pairs those names have to agree with.

Keep reading

Identity & Test Data Generator guides