Small territories and special codes are where a country data model finds out how much of its design was an assumption. The dataset works beautifully for the large, unambiguous, well documented places, and then a dependency, an overseas department or a reserved code arrives and something has to give.
The value of these cases is that they are cheap to add and unusually good at exposing structural problems. A model that handles them has usually made several important decisions explicit.
Why does the two letter code assumption break?
Because the code systems were designed for purposes that do not all require a two letter code for every place. The two letter system is the familiar one, but there is also a three letter system and a numeric system, and each exists to solve a different problem. The numeric system in particular arose to avoid the ambiguity of letters across languages.
Subdivision coding adds another layer. Some places appear only as subdivisions within a larger entity rather than as entities in their own right, and some appear in one system but not another. A model that treats the two letter code as the primary key of the world will simply not have a slot for those places.
The practical consequence is that the primary key should be something your system controls, not a value borrowed from a standard that has its own scope decisions. The standard code then becomes an attribute — important, searchable, but replaceable when the situation calls for it.
What should happen to reserved and retired codes?
Codes are added, reserved for future use, and withdrawn. Withdrawn codes do not disappear from the world; they remain in historical records, in archived documents and in old databases, and they will keep arriving in imports long after the code has been retired.
A system has to decide what to do with them. Rejecting them outright breaks the ability to read history. Accepting them silently as current invents a fact about the present. The workable middle is to accept them with a marker that says they are not current, and to keep that distinction visible wherever the value is displayed.
The same logic applies to codes that were reserved before they were ever used. A reserved code is not an error; it is a gap in the numbering that someone chose deliberately, and a validation rule that rejects it as malformed is wrong about the standard rather than about the input.
Former names and the names a place uses for itself
Names change, and the change is rarely simultaneous everywhere. For a period, a new name is official while the old name is still what most documents contain, and for place names the local form and the international form frequently differ.
Three kinds of name therefore have to be distinguished. The current official name, the name that appears in older records, and the name used locally by the people who live there. A system that keeps only one of the three will be wrong in one direction or another, and which direction depends on who is reading.
This is a place where cleaning data can destroy it. A well intentioned normalisation pass that rewrites every historical record to the current name makes the archive agree with the present and stops agreeing with itself.
Not applicable is not the same as unknown
Small territories make this distinction impossible to ignore, because they are the cases where a field genuinely does not apply.
A field that does not apply means the concept does not exist for that place, so there is nothing to look up and no correct value. A field that is unknown means the value exists and has not been established yet. These two produce the same blank cell and require opposite responses, and a checklist that records them as one state will keep producing work that cannot be finished.
Where the field is a postal code system, the difference is stark. The guide to postal code formats by country covers how much the field varies between places. Every country’s situation has to be classified before any validation touches it, because a system that assumes a field exists will mark every entry without one as incomplete, and a system that assumes it may not exist will accept genuinely bad input.
Unsupported is not the same as malformed
This is the distinction that most often gets lost, and it is the one with the clearest user-visible cost.
A value that a system does not support is not a value that is wrong. It may be perfectly well formed for its own place, in a format the system has simply never been taught to handle. Reporting it to the user as invalid tells them they have made a mistake when in fact the limitation is on the receiving side.
| What the system found | What it knows | What it should say |
|---|---|---|
| A value that follows the rules it implements | The value is valid | Accept it |
| A value that follows rules it does not implement | The value may well be valid | Cannot be checked here |
| A value that breaks rules it does implement | The value is invalid | Report the specific problem |
| A field that does not exist for this place | There is nothing to check | Not applicable, not missing |
Collapsing the middle two rows is the common failure. It produces a product that is confidently wrong at exactly the places where its knowledge is thinnest, and customers from those places are the ones who notice first.
For developers: make the third state expressible
A binary valid or invalid outcome cannot represent “not checked here”, so the validation result needs a third value that propagates honestly to the interface. That is a change to a type rather than a change to a rule, which is why it tends to be skipped under time pressure and regretted later.
Then test the boundary cases on purpose. Include a place whose code appears only in one system, a code that has been retired, a place with no postal code system, and a name that changed. These four cases are small, fast and unusually productive, and they are worth keeping in the suite permanently rather than adding when a ticket arrives.
The country and region directory is a reasonable place to see how a country and its subdivisions are displayed when they are treated as a first class entry rather than an exception, and the United States entry shows what a fully specified entry looks like next to a thinner one.
The territories, codes, names and field states used in this article are constructed boundary cases assembled to illustrate a design problem. They are not a dataset, they do not represent any real territory’s actual classification, and nothing here should be quoted as a fact about a specific place.
Next steps
Add the four boundary cases above to your suite and see which of them produce a wrong answer rather than an unhandled one. The guide to choosing countries for test data explains how to place such cases deliberately in a default set, and the guide to cross border address scenarios covers what happens when a place like these appears as one of several countries in a single transaction.