Tax ID validation is the practice of checking a tax identifier before it is stored, invoiced against or reported — and it is also the field where expectations are most often set too high. A well-designed check will stop transposed digits from entering your database. It will not tell you whether the business exists, whether the number was ever issued, or whether the company behind it is entitled to anything.
This article explains what a tax identifier is, why some formats can be checked arithmetically and others cannot, how to build validation in layers so that each layer answers a question it can actually answer, and why a failed check must never be reported as a missing company.
What is a tax identifier?
It is the identifier a tax administration assigns to a taxpayer. That is the whole definition, and it is deliberately broad, because countries organise tax registration in very different ways.
One country may have a single national tax number used for everything. Another may have separate numbers for different taxes, so a business holds one identifier for direct taxes and a different one for value added tax. Another may use a number that predates the tax system entirely, having been introduced for a different administrative purpose and later adopted for tax. Some jurisdictions use one identifier that doubles as the company registration number; others keep them strictly separate.
Two consequences follow immediately. First, “tax number” is not a single field in reality; it is a family of fields whose membership depends on the country. Second, any statement of the form “a tax number always contains x” is false somewhere, which is why your validation has to be written defensively.
Why do some tax numbers carry a check digit?
Because the issuing authority designed them that way. A check digit is an arithmetic redundancy: the last character, or one embedded in the middle, is computed from the others by a published rule, so a single mistyped or transposed character usually produces a value that fails the rule.
Where a country publishes such a rule, implementing it is genuinely useful. It catches the most common data-entry errors at the moment of entry, with no network call, no external dependency and no privacy question. It is cheap, fast and deterministic — and it is the strongest check you can perform locally.
Where a country publishes no such rule, the identifier is simply an allocation: the authority assigned it and recorded it, and no arithmetic can distinguish a real one from an invented one. This is more common than developers expect. Some very large jurisdictions issue plain sequential numbers with no redundancy at all.
The instruction that follows from this is blunt. Implement a check-digit algorithm only for the specific countries where you have verified that the algorithm is published and still current. Do not guess an algorithm from the shape of the number, and do not assume that because one country’s format looks similar to another’s, the same rule applies.
| What the number looks like | What you can honestly check |
|---|---|
| Published algorithm with a check digit | Character set, length, and the arithmetic rule |
| Plain allocation, no published rule | Character set and a generous length bound only |
| Format with a country or authority prefix | That the prefix matches the country claimed |
| Format that also serves as a registration number | Whatever the register itself can confirm |
Does passing validation mean the number is real?
No. Format validity and existence are different properties, and conflating them is the most common design error in this area.
A syntactically perfect number can belong to nobody. Anyone who understands the pattern can produce one that satisfies every local rule, including the check digit, because the check digit is designed to catch accidents rather than adversaries. A number can also be entirely real and fail your local rule, if the rule was written against an older layout, a different country, or a document that formatted the number unusually.
Existence is established only by the authority that issued the identifier. Where a tax administration publishes an official enquiry facility, that facility is the answer to the existence question; where it does not, existence may be confirmable only from documents the business supplies. That is a genuinely weaker assurance, and it should be described as such rather than dressed up as verification.
There is also a subtlety that surprises teams: even a confirmed registration is a point-in-time fact. A number can be valid today and cancelled next month, so an aggressive cache of “this number is valid” eventually conflicts with reality. Cache deliberately, and prefer caching the answer with a timestamp over caching it forever.
How should validation be layered?
In three passes, ordered by cost, with the cheapest and most certain first.
The first pass is local and structural: is the field non-empty, is it within a generous maximum length, does it contain only characters that appear in some format, and does the country claim match any prefix present. This pass should almost never reject a plausible input; its job is to catch empty and obviously damaged values.
The second pass is country-specific: apply the published check-digit rule where you have one, and only there. Report the outcome as a typo warning attached to the field, phrased as “this does not look right, please check it”, because that is what a check-digit failure actually means.
The third pass is the authoritative lookup, asynchronous and never blocking the form. It has three outcomes worth distinguishing — confirmed, not found, and unavailable — and it should be able to say which of them occurred. A lookup that cannot tell “no such number” from “the service is down” will eventually reject a legitimate customer on a bad network day.
For developers: errors, caching and audit trails
The design work is mostly about honesty in error handling. Each layer should produce a distinguishable result, and the message the user sees should name the layer that failed. “We could not reach the tax authority, your details are saved and we will check again” is a usable message. “Invalid tax number” attached to a value the user copied from an official certificate is not.
Consider an explicit result type rather than a boolean. Four states — unverified, shape valid, lookup confirmed, lookup not found — cover the ground without forcing a distinction the system cannot make. A boolean collapses them and guarantees that some legitimate value is eventually treated as false.
Caching should be asymmetric. A confirmed result is worth holding for a reasonable period; a not-found result should expire quickly, because registrations change and yesterday’s absence is a poor reason to reject today’s customer. An unavailable result should not be cached as a negative at all.
Audit what you checked without turning the log into a copy of the customer database. Record the fact and time of a check, the layer applied, and the outcome; avoid duplicating the raw identifier into systems that do not need it. Where the identifier must be retained, treat it as business data with the same care as any other — and keep generated values clearly distinct from real ones, which is what the synthetic tax identifier fields on this site are for. The VAT number formats article covers the most common member of this family, and test company data shows what the rest of a synthetic business record looks like.
Next steps
Inventory every tax-number check in your product and label each one with the layer it implements. Any check that rejects a value on the basis of a guessed pattern should be demoted to a warning, and any message that says “does not exist” on the strength of an arithmetic test should be rewritten to say what was actually tested. Then use the company data generator to produce records for several countries and confirm your form handles a country it knows nothing about without refusing the entry outright.