Menu

International Address Forms: A Per-Country Test Checklist

How address line order, postal code patterns, phone prefixes and CJK concatenation differ by country, and where international forms break.

Published

  • forms
  • internationalization
  • addresses
  • postal-codes

An address form that works for one country is usually a form written for the country its author lives in. Add a second country and the assumptions begin to collide: how many lines an address needs, which field carries the postal code, whether a region is required at all, and which characters a user is allowed to type. The failures are rarely crashes. They are addresses that save without complaint and cannot be delivered to.

This guide is a checklist of what actually differs around the world and where forms usually break on each point. It is written for the person who has to produce the test cases, and every item on it is something a reviewer can verify before a release.

What changes when a second country is added

Three assumptions break at once, and they tend to break in the same order for almost every form.

The first is shape. A template built around address line one, address line two, city, state and postal code encodes a line count and a field order that not every country uses. The second is obligation: fields marked required for everyone are wrong for the countries that do not use them. The third is content: a rule written against one country’s format rejects legitimate values from another, and rejection is at least visible, which is more than can be said for the alternatives.

Silent breakage is the category worth testing hardest. A truncated postal code, an uppercased code that was case sensitive, or a city placed in the wrong slot all produce a record that saves cleanly and fails much later, in a carrier system or a tax calculation nobody in the form team can see. The guide to address validation and normalisation is where that cleanup work belongs.

Address line order and line count

Western addresses run from smallest to largest: name, street, city, region, postal code, country. Japanese, Chinese and Korean addresses run the other way, from country down through prefecture or province, city and ward to the street and building. A form that renders a fixed Western template and squeezes a Japanese address into it produces something that reads wrong to a human even when every field is populated.

Line count varies as well. A British address commonly needs four or five lines including the post town, a US address usually fits in three, and a Japanese address often needs two lines for the block and building detail alone. Two general address lines is a defensible minimum, one is not, and four is generous rather than excessive.

The label for the middle layer is not the same word everywhere either. It is a province, a region, a prefecture, an oblast, a governorate, an emirate, a department or a county depending on where you are. The label matters less than allowing the field to be optional for countries that do not use one in postal addressing. The international address format article covers the field names themselves.

Where does the postal code sit, and what shape is it?

Both the position and the pattern differ, and a rule written for one country is the classic defect in this area.

Country Position in the rendered address Shape
Germany Before the city Digits only
France Before the city Digits only
Japan Before the city Digits with a hyphen inside
United States After the city and region Digits, with an optional hyphenated extension
United Kingdom After the post town Letters and digits, with a meaningful internal space
Canada After the province Alternating letter and digit
Netherlands After the city Digits followed by letters
Ireland After the town An alphanumeric code of its own kind

Two further cases defeat pattern matching entirely. The first is a country whose code is genuinely alphanumeric, so a digits-only rule refuses valid input. The second is a country with no postal code system in everyday use, where the field has to be optional or hidden rather than required, because a required field with no valid answer forces users to invent one.

Length limits are the same bug pointed the other way. A code that carries an internal space is longer than the digits or letters alone, so an input capped at five or six characters truncates a value that was perfectly correct when typed. Truncation is worse than rejection, because the form reports success and the damage is found by someone else.

Are postal codes a pattern or a geography?

They are geography, and the pattern is only the surface of it. A postal code identifies a delivery area, and its leading characters carry location in most systems.

That has a direct testing consequence. A code can satisfy the shape rule and still be impossible for the city beside it: a code whose leading characters belong to one part of a country cannot be right for a city in another, and both pass a regular expression without complaint. Pattern validation catches typos. The relationship between code and region requires reference data, and it is the check that catches a genuinely wrong address rather than a mistyped one. The per-country shapes are catalogued in postal code formats by country.

If you can only test one of the two, test the relationship. A pattern failure is visible to the user and gets fixed; a pattern-valid code in the wrong place is invisible until a carrier or a tax engine looks the address up.

Phone numbers, trunk prefixes and digit counts

Address forms usually collect a phone number, and it behaves per country for two reasons.

The national trunk prefix is the leading digit used for domestic dialling, and international format omits it. A form that stores what the user typed and prepends a country code produces a number that cannot be dialled. Accept national input, strip the trunk prefix where present, and store the result in international format. The same principle is worked through in phone prefix and locality matching.

Length is not uniform either, so validation should consult the country record instead of assuming one global rule. The country code is a third variable: a selector that changes the dialling code without changing the expected national length will happily accept a number of the wrong size for the country chosen. Keep the code derived from the selected country and validate the national number on its own.

What breaks when Latin and CJK addresses share a form?

This is where most international forms quietly give up. CJK addresses are concatenated without spaces between the administrative layers and written largest to smallest; the romanised version is smallest to largest and gains spaces. A form that joins fields with a comma and a space and prints the result on one line gets both versions wrong.

Character width adds a second failure. Digits and Latin letters typed on a Japanese or Chinese keyboard can arrive at full width, which fails a numeric check even though the value is right. Normalising the text before validating solves it, and the order matters: normalise first, then check.

Name order is the third. Names written family-name-first break any routine that splits a value into a given name and a family name and rejoins it in the opposite order, and a name that does not split into two parts at all cannot be represented by a pair of fields in the first place.

Which input classes break naive validators?

Three, and none of them is exotic.

  • Full-width digits and letters, which look almost identical on screen and are not recognised as digits by a parser.
  • Diacritics and special letters, from the German sharp s to Nordic vowels and Turkish dotless letters, which count differently in a character limit than in a byte-limited column and which a well-meaning romanisation rewrites without asking.
  • Honorifics and multi-part names, which belong to the address block rather than being decoration and which a two-field name split cannot hold.

Test all three with real strings in the real script rather than a transliteration. A form that handles a romanised name is not evidence that it handles the characters a user will actually type.

A twenty-minute manual pass across three countries

Automated checks catch regressions; a manual pass catches assumptions. Three countries that disagree on nearly every axis will surface more in twenty minutes than a week of unit tests written by whoever built the form.

  • United States: postal code after the city, a two-letter region, a digits-only code with an optional hyphenated extension, and a ten-digit national phone number. Confirm that the second address line is honestly optional.
  • Germany: postal code before the city, digits only, and a federal state most users never type because the code already implies it. Confirm that a required state field is not a dead end for foreign residents.
  • Brazil: a code written with a hyphen, a two-letter state abbreviation, and a neighbourhood field that many templates do not have. Confirm that the neighbourhood has somewhere to go.

Then repeat five operations for each country: enter a valid sample and confirm it saves; enter a postal code from a different region and see whether anything objects; paste a whole address into line one and watch what happens; switch the country selector after typing and observe what the form clears; reload the saved record and compare it with what was entered. Run the whole pass once at phone width, because half of these failures only appear when the layout collapses. The checkout address form test cases article extends this pass to the rest of the checkout flow.

Where these forms usually break

  • A region field required for every country, including countries that do not use one.
  • A postal code pattern hard-coded to one country’s shape, plus a maximum length that truncates longer valid codes.
  • Case transforms that rewrite a code which is only valid as entered.
  • A country selector placed after the address fields, forcing the form to be filled in the wrong order.
  • A second address line labelled as an apartment field when it is the only place a CJK building number can go.
  • Placeholders used as labels, which disappear the moment typing starts and take the format hint with them.
  • No room for long division names, which are ordinary in many countries and rare in the ones a form is usually tested with.

Next steps

Sort the checklist by risk rather than by field. Every release, test one valid and one invalid sample per supported country, confirm that requiredness follows the country, push every free-text field to its longest realistic value, and do a save-and-reload round trip. Then assert the layout as well as the validator, because field order, label text, optionality and maximum lengths are what the user actually meets.

Generate the samples rather than typing them by hand. The address generator returns the same record for the same key, which keeps a snapshot stable across runs, and the country directory tells you which divisions and cities a sample for a given country should use. Hand-typed samples drift, and the drift stays invisible until a form change breaks them.

Every value described on this page is invented for software testing. None of the samples represents a real address, and the country rules quoted here describe formats rather than anyone’s residence.

Keep reading

Fake Address Generator guides