Menu

Address Validation vs Normalization: Two Different Jobs

Address validation and normalization get used as synonyms, but one reformats and the other tries to prove existence. Knowing which you need changes the design.

Published

  • test data
  • address
  • validation

Address validation and normalization are treated as one activity in most projects, and the confusion costs teams real time. One of them reformats a value so that two spellings of the same address match; the other tries to answer whether the address exists and can be delivered to. They use different data, fail in different ways, and — in a test suite — deserve entirely different assertions.

This article separates the two, explains why no authority holds a complete global address list, and describes what a system can honestly claim after running either one.

What normalization actually does

Normalization is a rewrite. It takes what a user typed and produces a canonical version of it: consistent casing, expanded or abbreviated street types according to one convention, standardised directional words, a consistent separator, trimmed whitespace, and a postal code written in one agreed form. Nothing is verified. A string that never corresponded to any real location normalises just as neatly as a correct one.

Its value is comparison. Two records that describe the same place but were typed differently become byte-identical after normalisation, so deduplication, matching and search start working. It also stabilises storage: one form in the database means downstream consumers stop reimplementing the rules.

Normalisation is cheap, deterministic and safe to run anywhere. That combination is why it belongs at the earliest boundary the data crosses — the form handler, the import script, the API entry point — and only once. Normalising repeatedly in different layers is how a value ends up oscillating between two spellings.

What does validation actually try to prove?

Validation asks whether a value is acceptable, and the honest answer depends entirely on how much evidence you have. Three different questions hide behind the single word.

The first is structural: do the required fields exist, are the characters in the allowed set, is the length plausible. This needs no reference data at all and can be done anywhere.

The second is relational: does the postal code fall in a range used by the stated region, is the state code consistent with it, does the country agree with its subdivision code. This needs reference data but not an address database, and it catches a large share of genuine typos.

The third is existential: is there actually a building at this number on this street. This needs an authoritative local dataset, and only some countries publish one that is complete enough to answer it.

Most systems claim the third and implement the second. That is not necessarily wrong, but it should be a conscious decision rather than an accident, because it determines what the error message can honestly say.

Question What it asks What it needs
Structural Whether the required fields exist, the characters are in the allowed set and the length is plausible No reference data at all
Relational Whether the postal code falls in a range used by the stated region, the state code is consistent with it and the country agrees with its subdivision code Reference data, but not an address database
Existential Whether there is actually a building at this number on this street An authoritative local dataset, which only some countries publish

Is there such a thing as a global address database?

Not one that is complete, authoritative and current. Postal operators maintain their own delivery data for their own territories, under their own terms, and many of them do not publish an address-level list at all. Where national datasets exist, they cover that country, in that country’s language and structure. There is no single register that a vendor can consult to decide whether an arbitrary address anywhere on earth exists.

What vendors actually do, in various combinations: use their own aggregated reference data, match against postal code and administrative geographies, apply plausibility rules, and — for some countries — use a licensed local dataset. Every one of those is partial. A service that reports an address as valid is usually reporting that it looks consistent with what the service knows, and a service that reports it as invalid may simply be missing coverage.

The practical takeaway is that “address verified” is not a claim your system can support globally. “Structurally valid and consistent with the reference data available to us” is a claim it can support, and it is the one worth putting in the documentation.

Why order matters: normalise, then check

Run normalisation first. Reference checks are written against a canonical form, so an unnormalised value will fail them for reasons that have nothing to do with correctness: a postal code with the separator in a different place, a street type abbreviated differently, a case difference in a region name.

The reverse order is worse than useless. If the reference check runs first and rejects the value, you lose the chance to fix a string that was merely untidy, and the user sees an error they cannot act on. If it runs first and accepts, you still need the normalisation afterwards, so nothing is saved.

A second ordering rule applies to error handling. Distinguish “this could not be parsed” from “this conflicts with reference data” from “this is not in our coverage”. The three need different messages and different follow-up actions, and collapsing them into one failure is what makes address validation feel arbitrary to users.

For developers: rules, overrides and fixtures

Do not try to validate an address with one regular expression. The variation lives in the parts a regular expression is worst at — street names, unit details, non-Latin scripts — and the certainty lives in the parts you can check against a list. Reserve patterns for the length and character checks, and keep the reference checks in data.

Always allow a manual override. Reference data is incomplete and sometimes wrong, and a genuinely correct address that your checker does not recognise must remain liveable. Log the override with the value that was rejected, so you can tell whether the checker is improving or simply annoying people.

Keep normalisation and validation as separate steps in your own code, with separate names. A function called address-check that sometimes rewrites its input and sometimes refuses it is the kind of thing that makes bugs untraceable. The same separation is what makes it possible to run normalisation over historical records while leaving validation as an input gate, a distinction familiar to anyone building test fixture address data.

For a test suite, assert the two behaviours independently: that a given messy input normalises to an expected canonical string, and that a given inconsistent pair of fields is rejected. When the reference list changes, only the second set of tests should move.

Next steps

Write down, in one sentence, which of the three questions your system actually answers today, and check whether your error messages claim more than that. Then add a fixture for a correct address your checker rejects, and decide what should happen to it. The address generator produces values that are structurally valid by construction, which makes it a convenient positive baseline alongside the deliberately broken samples you keep. Structurally valid is all they are: those values cannot be delivered to a destination and they prove nothing about residence. For a walkthrough of one inconsistency that is easy to detect, see checkout address form test cases.

Keep reading

Fake Address Generator guides