An international address format is not one format. It is a family of conventions that disagree about what comes first, what a division is called, and whether the recipient’s name sits above or below the building. Anyone building a form, a label or a printed envelope meets those disagreements quickly, and the cheapest fix — one template with a couple of optional lines — stops working as soon as a second country joins the list.
This article covers how addresses are actually ordered around the world, which fields vary, why the script matters, and how to model the data so that display decisions stay separate from storage decisions.
Two directions of writing an address
Addresses are written from the smallest unit outward in some places and from the largest inward in others, and the two habits produce mirrored line orders.
In much of East Asia, an address traditionally runs from the large unit to the small: country or province first, then city or ward, then district, then street, then building number, then the unit. Western conventions run the other way: house number and street, then city, then region, then postal code, then country. Neither order is more logical; each groups information the way the local postal service and the local reader expect it.
International mail adds one more convention on top. Because the destination country is what sorting machinery needs first, it is traditionally placed as the final line, alone, in capitals or otherwise prominently. That is a routing convention rather than a data rule, and it is why printed envelopes often look nothing like the order in which a form collects the fields.
| Convention | Where it is the habit | Line order |
|---|---|---|
| Large unit outward | Much of East Asia | Country or province, then city or ward, then district, then street, then building number, then the unit |
| Small unit outward | Western conventions | House number and street, then city, then region, then postal code, then country |
| Routing convention | International mail | The destination country last, alone, in capitals or otherwise prominent |
Can one template handle every country?
Only badly. A template hard-codes an assumption about which fields exist, which are on the same line, and what order they appear in, and those assumptions are exactly what changes between countries.
A single fixed template will place a house number in front of a street name in a country that puts it after, insert a comma where the local convention has a marker character, or reserve a line for a state that the address never uses. Consider what happens in practice: some countries habitually omit a division entirely, locating an address by town and postal code instead, so a mandatory region field forces users to repeat the city or pick an approximation. Others use a division level that has no equivalent in the template, and the name gets dumped into whichever field is free.
The fix is smaller than it sounds: keep the data complete and let the display layer decide line order. The ordering belongs to the presentation, and no field should depend on it.
What is the difference between a state, a province and a prefecture?
They are all first-level administrative divisions, and the word for them differs by country — state, province, prefecture, region, canton, governorate, emirate, oblast. The label matters to the reader and not at all to the data model, which is why calling every one of them “State” creates more trouble than it solves.
The problems are concrete. A user in a country that calls the division something else has to guess what the form means. A support script that filters records by state excludes records whose division lives under a different name in the source system. A validation rule checked against a list of US states rejects nearly every address outside the United States.
Keep the field generic in the schema and localise the label at render time. Store the value together with its country code, since a division name is only meaningful alongside the country it belongs to; two countries can share a division name and mean different places. The family resemblance among these codes is described in the guide to ISO country and subdivision codes.
Script, transliteration and the quiet failure of Latin-only fields
Many addresses are written in a script other than Latin. Some destination countries require the local script for domestic delivery and accept a romanised version for international routing; others expect the romanised form on inbound mail. Neither arrangement is universal.
Two things go wrong when a system assumes Latin characters. The first is outright rejection: a character set that excludes non-Latin letters turns a valid address into an error, and the user cannot fix it because the address is correct. The second is quieter — a field that accepts the characters but normalises, transliterates or truncates them, so the stored value no longer matches what the user typed and a later comparison fails.
Characters are not the only hazard. Some scripts are written without spaces between the address components, so a parser that splits on whitespace finds one enormous token instead of six fields. Others use a marker between the street name and the number where Latin conventions use a space or a comma. Length is not uniform either: a romanised address is usually longer than its original, so a field sized to the local script can be too small for its own transliteration.
Dealing with building names, unit numbers and districts
Real addresses carry units that templates rarely anticipate: apartments, floors, blocks, towers, entrance numbers, building names, sub-district names, and directions in the form of a landmark. The local list of these identifiers differs as much as the division names do.
Two rules keep them manageable. Collect the unit details in their own field rather than appending them to the street line, so the street survives for parsing and geocoding. And when a country’s addresses commonly include a level that your template lacks — a district beneath the city, a building name above the street — add a field rather than concatenating. A field whose meaning is one thing is cheap; a field that means “whatever else there was” is where data quality goes to die.
For developers: one field, one meaning
The design principle that survives contact with the most countries is the least clever one: give every piece of address information its own field with one meaning, and never let a field’s name imply a country.
Practically, that means a country code field, a division field that is not called state, separate fields for district, city, street, house number and unit details, a postal code field, and a free-form line for information that does not fit any of them. Keep the fields optional in the schema and mandate what you need per country at validation time, where the country is known.
Add a per-country display definition, kept apart from the data: the order of lines, which fields are joined, whether the division appears in the address block, and where the postal code goes. That separation lets you add a country by editing one definition rather than reopening the form logic. It also means a bug in layout cannot corrupt stored data.
Last, be careful with concatenation. If you build a single-line address for an API call, make the separator part of the display definition too, because joining with a comma everywhere produces an address that reads correctly only in the countries that use commas.
Next steps
Pick three countries you actually serve and write out, by hand, how each of them orders the same information. Where the three disagree, you have found the seams your template is going to break on. Then generate one address per country in the address generator and paste each into your form to see whether the layout still makes sense; the guide to address data in test fixtures explains how to keep those samples reproducible. Keep in mind what those samples are: generated values shaped like real addresses, intended for testing rather than delivery, and no indication of where anyone actually lives.