Menu

Country versus language: three layers of localization

Country versus language is a distinction that decides how a test matrix is built. Separate the interface language, the content region and the data format.

Published

  • localization
  • test matrix
  • country data

Country versus language is one of those distinctions that everybody agrees with in the abstract and violates in the first fixture. A test matrix is written as one row per language, a country is attached to each row to make the data look realistic, and the two axes quietly become one.

This article pulls them apart. It describes the three layers that the word localization usually hides, why a language tag and a region code do different jobs, and how to sample two axes without pretending they are the same axis.

Why can country and language not be derived from each other?

One language can be an official language in many countries, and one country can have several official languages. Both directions of the relationship are many, so neither value determines the other.

The consequence for testing is immediate. If every language is paired with exactly one country, the matrix has a diagonal shape: it contains the pairs that someone found natural and none of the others. The interesting failures live off that diagonal — the same language rendering a different format, and the same format appearing under a different language.

A second consequence is subtler. Because the diagonal pairs tend to be culturally familiar to whoever wrote the matrix, they are also the pairs where the team’s assumptions are most likely to be correct. The matrix is strongest exactly where it is least needed.

Localization is really three layers

Treat localization as three decisions that happen to share a word.

Layer The question it answers Typical owner
Interface language Which language are the labels, buttons and messages in? Content or product
Content region Which market’s rules, prices and offerings apply? Business
Data format Which conventions govern dates, numbers, names and addresses? Data or engineering

Each layer can be set independently, and in real systems they frequently are. A reader may browse in one language, be served the rules of a market they are travelling in, and enter an address that follows the conventions of a third place.

Once the three layers are separated, most localization defects become describable. A defect is a case where two layers were assumed to move together and they did not.

The language tag and the region code do different jobs

A language tag describes text. It says which language a string is written in and, sometimes, which script or variant. A region code describes the conventions a value follows: how a date is ordered, how a number is punctuated, how a name is arranged.

They are not interchangeable, and one is not a more precise version of the other. A system that stores only a language tag has thrown away the information needed to interpret a date, and a system that stores only a region code has thrown away the information needed to choose a message catalogue.

Keeping both is not redundancy; it is the minimum required to describe a rendered page. The bug to watch for is the shortcut in the other direction — using the language tag as the switch for data formatting, which works right up until the first country that shares a language with another.

What breaks when a format follows the language?

The classic failure is a value validated against the conventions of the wrong country because the two share a language. The user’s language is set correctly, the format rule is chosen from that language, and the value is rejected even though it is well formed for the country the user actually lives in.

There are quieter versions. A name field that reorders parts because the interface is in one language rather than because the record belongs to a region with that ordering. A numeric field that accepts a decimal separator from the wrong convention and stores a value that is off by orders of magnitude. A date that is interpreted day-first in one place and month-first in another and lands in a report as a plausible but wrong instant.

None of these are language problems. They are all cases of a region decision being made by a language input, and the fix is structural rather than a better lookup table.

How should the test matrix sample two axes?

Sample each axis on its own terms, then cross a small number of deliberately chosen combinations rather than every cell.

Start with the language axis and choose entries that differ in what the interface needs: a language that expands text considerably, one that needs a different script, one whose sorting rules differ from the Latin default. Then take the region axis separately and choose entries that differ in what the data needs: a region where a field does not exist, one where a value is unusually long, one whose conventions conflict with a neighbour sharing its language.

The crosses that matter most are the off-diagonal ones — the same language in two regions, and two languages in one region. Those four cells catch the class of defect that “one language, one country” cannot express, and they are cheap to add once the axes are stored separately.

For the linguistic detail of how names and text behave per locale, there is a separate guide on name data by locale, which stays with the language axis; the region axis is what this article is about.

For developers: let the region drive the format

Store both values, and be explicit about which one drives which decision. Format rules should be selected from the region; message catalogues and text layout should be selected from the language; and neither should be inferred from the other at render time.

Keep the two settings in different places in your configuration, with names that say what they are. A field called locale that holds a language tag is a permanent invitation for the next developer to use it as a region. A field that holds a region code should never be described as the language a user speaks.

Then assert the separation in tests. A test that renders one language across several regions, and one region across several languages, fails loudly the moment someone reintroduces the shortcut. The country and region directory and a country page such as the Japan entry show how a single country’s conventions are described when they are kept next to each other, which is a more reliable reference than a language name.

Nothing here should be read as a description of real traffic. The language and region combinations used as examples in this article are constructed sampling choices, not observations of any real user base, and they say nothing about where anyone lives or what anyone speaks.

Next steps

Write down the three layers for one screen your team owns, and name the value that feeds each of them. If two layers are fed by the same value, you have found the shortcut. The guide to testing country select fields takes the region axis down to the control the user actually touches, and the guide to region groupings and market tiers covers how regions themselves get defined.

Keep reading

Address & Identity Data Formats for 86 Countries guides