Address format by country is not a matter of translation. The order of the elements, the presence of a postal code, the number of administrative layers and the way a house number relates to a street all change when you cross a border, and a form designed around one country’s convention quietly rejects valid input from most of the others.
This guide sets out the main structural differences, explains why the postal code is the field most often misdesigned, examines the compromise that international forms use and what it costs, and offers a way to model a multi-country address that can actually be tested. By the end you should be able to look at an address form and predict which countries will fail in it.
How does the order of elements differ?
The two dominant conventions run in opposite directions. Anglo-American addresses are written smallest to largest, with the recipient’s name, the house number and street, the locality, the administrative area and the postal code in that order. Japanese addresses are written largest to smallest, beginning with the prefecture, then the municipality, then the district or neighbourhood, then the block, then the building and finally the apartment.
Turkey and several of its neighbours follow the largest-to-smallest direction as well, with the province leading and the neighbourhood and street following. Latin American addresses often sit between the two, listing the street, then the number, then the district or colony, then the city and the administrative division, with the postal code either before or after the city depending on the country.
German-speaking countries usually put the postal code and the city on the same line with the postal code first, which is the reverse of the American arrangement. French addresses place the postal code before the city as well, and Dutch addresses do too. When a form labels a single field as city and expects the postal code to be separate, the display order it produces will be wrong for all of those countries even though the underlying data is correct.
The practical consequence is that address rendering and address storage should be separate concerns. Store the elements in their own fields, and assemble the display string according to the country’s convention at the point of display. A system that stores a pre-rendered single line has already decided the order and cannot change its mind for a different destination.
What does the postal code look like in different countries
The postal code is the field where the widest variation sits, and it is the one most often modelled as a fixed pattern. The United States and Turkey use five digits. Germany, France, Spain, Italy and Mexico use five digits as well, but the leading zeros are significant and a numeric database column will silently drop them. The United Kingdom uses an alphanumeric format of variable length with a space in the middle, and both the space and the letter case matter for matching.
Canada uses an alphanumeric format that alternates letters and digits. The Netherlands uses four digits followed by two letters. Poland uses five digits with a hyphen after the second. Japan uses seven digits, usually written with a hyphen after the first three, and Brazil uses eight digits written with a hyphen after the first five. Ireland has one of the few genuinely alphanumeric mixed systems with a routing key.
Then there are the countries with no postal code at all. Several jurisdictions do not operate a nationwide system, and some have introduced one only recently and incompletely. A form that marks the field as required will reject every address from those countries, which is a failure that a test suite built only from countries with postal codes will never discover. The postal code formats by country article collects the patterns in one place.
The storage rule that follows from all of this is simple: keep the postal code as text, never as a number. Leading zeros, letters and internal separators all break a numeric column, and the loss is silent. The same rule applies to house numbers and building numbers in countries where those values carry letters or slashes, because a field that worked for one country’s addresses will quietly corrupt another’s.
Which fields exist in some countries and not others
The field set itself is country-dependent, which is a harder problem than order or format. Some countries have a state or province level; some have several administrative layers; some have a neighbourhood or district level that is essential for delivery; and some have none of those.
The administrative level that most international forms omit is the one that carries the most weight in delivery. In Turkey it is the district and neighbourhood. In Japan it is the municipality and the block. In many Latin American countries it is the colony or barrio. In Indonesia the address carries a village and a district beneath the city. An international form with city, state and postal code cannot express those addresses, and the users affected will either put the missing element into the street line, where it corrupts the data, or abandon the form.
Conversely, some fields are mandatory in one country and meaningless in another. A state is a required field in the United States, Australia and Canada, an optional field in many others, and a concept that does not exist in a further group. If your template marks it required for all countries, every address from a country without states carries a fabricated value in that column, and any later analytics on state-level distribution become fiction.
The honest design keeps a single form with a per-country field configuration. The configuration says which fields are shown, which are required, how they are labelled, what pattern the postal code follows, and what values the administrative field may take. That configuration is data, and like any data it should be tested country by country rather than assumed.
What is the address line compromise and what does it cost?
The address line compromise is the pattern in which every address in the world is flattened into two or three generic lines plus a city, a region and a postal code. It exists because it was the cheapest way to make an international form work, and it persists because changing it is expensive.
Its strength is that it never blocks an address. Anything can be typed into a free-form line, so users are never rejected for entering a format the form did not anticipate. Its weakness is that it stores an unstructured string where a structured value was needed, so anything downstream that groups, validates, routes or deduplicates by geography has to parse text. There is also a versioning problem: as countries revise their administrative divisions, the flattening hides which convention a historical row was written under, and a change of naming in one province becomes indistinguishable from a typo.
The cost appears later and in specific ways. Address normalisation becomes unreliable because the parser has to guess which token is the city. Deduplication fails because the same address written twice in different orders looks like two addresses. Tax and delivery logic that depends on a postal code or a region cannot run at all when the value is buried in a text line.
There is a second cost that is easy to miss: the compromise hides its failures. A form with a free-form line accepts an address with no postal code from a country that requires one, and accepts a postal code from the wrong country, and reports success. The international address format article discusses the pattern in more depth, and the cross-border address scenarios article looks at what happens when the same person has addresses in two countries.
How should a multi-country address model be designed
Start with a country field that drives everything else, and treat every other field as conditional on it. The country determines which fields appear, which are required, what pattern the postal code takes, and what list the administrative field draws from. Nothing downstream should be hard-coded to a single country’s rules.
Keep the postal code as text with a normalisation step rather than a validation-only step. Normalising means removing the spaces and separators you do not need, uppercasing where the country’s convention is uppercase, and preserving leading zeros. Validating means checking the pattern for that country. Doing both in that order avoids rejecting input that is correct but written differently.
Model the administrative hierarchy as nested data rather than a flat list. The locality belongs to a region, and in many countries to a further layer between them. Storing the relationships rather than a single level means a new country can be added by adding rows rather than by changing code, and it makes the consistency checks expressible as data.
Then test the model by country, not by field. For each supported country, generate an address, assert that it passes your own validator, assert that the region and postal code agree, and assert that the rendered display order matches the country’s convention. A suite that tests the fields individually will pass while the combinations are wrong, which is the same failure mode that affects hand-built address data everywhere.
Finally, keep a record of why each country is configured the way it is. A postal code pattern or a field requirement is a statement about a country at a point in time, and the next person to touch the configuration needs to know which authority the rule came from. The country data coverage checklist article turns that into a reviewable list.
Telephone and address belong in the same test. A record where the number carries one country’s calling code and the address sits in another is a consistency defect that a country-aware model can detect, and the phone prefix and locality matching article explains the relationship. The countries tool on this site exposes the per-country field sets and conventions, and the country coverage checklist article describes what a maintained dataset should contain for each one.
What does a country-aware test dataset require
It requires per-country rules rather than one global rule, and it requires those rules to be maintained. Postal authorities change formats, administrative divisions are created and merged, and postal codes are reallocated. A dataset that was correct when it was assembled and has never been refreshed will eventually disagree with the authority it came from.
It also requires provenance. Knowing which source each country’s rules came from is what makes it possible to refresh them without guessing, and the country data freshness and sources article sets out the sources worth tracking. A dataset without provenance can only be replaced wholesale, which is why many are never refreshed at all.
Finally, it requires a clear boundary in the data itself. Addresses generated for testing are synthetic records with a correct structure and correct internal relationships, and they resolve to no real premises. They must not be used for real delivery, to prove a residence or an identity, to open or register accounts, or to get past any verification step that asks where someone actually lives.