Ask ten countries what a national identity number looks like and you will get ten different answers, not because the numbers are secret but because each register was designed on its own. Some states publish the arithmetic that closes the number, some publish only a description of its shape, and some publish nothing beyond the fact that the number exists.
That unevenness is the defining feature of this field, and it is the reason a generic international validator cannot be written. What can be written is a validator that knows, for each country it supports, how much it is entitled to check — and that says so plainly in the verdict.
Why can national ID numbers not be validated the same way?
Because there is no international authority that issues national identity numbers. Each one belongs to a national register with its own legislation, its own conventions and its own history, and nothing obliges a country to document its numbering for outsiders.
Where documentation exists, it also varies in kind. A published check-digit algorithm is a complete specification: given a candidate string, the arithmetic can decide. A published format is a partial specification: it can reject a malformed string but can never approve a well-formed one as genuine. A register that publishes neither leaves the outside world with nothing to implement, whatever the number’s internal logic may be.
Designing around that variety means treating coverage as a matrix rather than a boolean. For each country your product touches, record which of the three levels applies, and let the verdict follow from the level instead of pretending every country offers the same guarantee.
All identity values discussed here are illustrative descriptions of published rules rather than reproductions. No real identity number appears in this article, and a rule being described does not imply that any particular string was ever issued.
Countries that publish their check digit algorithm
Some registers specify the closing character completely: the weights, the modulus and the rule for the final digit are all in the public documentation. Those are the easiest cases, and they are also where a validator adds the most value, because a single mistyped character is caught before the number reaches any register.
The family of arithmetic involved is the usual modulus-based one, surveyed in check digit algorithms, and a national scheme is usually a specific parameter set within it. A handful of countries apply the doubling-and-summing member of the mod-10 family; others use a mod-11 arrangement, which is why the remainder-to-character mapping differs from case to case.
Even here, precision matters. Knowing that a country uses a modular check is not the same as knowing its parameters, and implementing a plausible approximation produces a routine that accepts most values and quietly disagrees with the standard on the rest.
Countries that publish a format but no algorithm
The middle tier is the largest and the most awkward. The register documents how many characters a number contains, which of them are digits, and often which component identifies a region or a birth period, but it never publishes an arithmetic rule. The United States social security number and India’s permanent account number both sit in this tier — widely used, thoroughly documented in shape, and without any public check-digit algorithm.
A validator facing a format-only scheme can still do useful work. It can reject strings of the wrong size, strings containing illegal characters and strings that contradict the documented component structure. What it cannot do is compute anything from the body and compare, because there is nothing published to compute.
The verdict must reflect that. A format-only scheme produces a format-only result, and the wording should say that the shape is confirmed while validity cannot be. Reporting such a value as valid is a false statement of exactly the kind that the four-verdict vocabulary exists to prevent.
What happens when one number matches several countries?
Two countries can agree by accident. If both use a similar length and a similar character set, a single string can satisfy both sets of rules completely — and if the two schemes also use comparable arithmetic, the same string can even pass both check digits. Nothing has gone wrong; overlapping specifications simply produce overlapping membership.
The temptation is to guess. A form that silently assumes the most likely country, or the country of the user’s browser locale, invents information that the input did not contain. Inferring a nationality from a string is both unreliable and a categorically different activity from validating a format.
The honest presentation is a list. Show every scheme the string satisfies, in a stable order, and let the user — or the downstream system with better context — choose. When nothing matches, report an unrecognised scheme rather than an invalid number, because the second statement is one the validator is not in a position to make.
Length and character set are only a starting point
Both checks are cheap, and both are frequently mistaken for the whole job. Length alone catches dropped and repeated characters; the character set alone catches the wrong alphabet. Neither says anything about the relationship between the characters.
That relationship is precisely what a genuine national number has and a random string does not. Where the register publishes the relationship, a validator can test it; where the register keeps it private, no amount of cleverness recovers it. Attempting to reverse-engineer an unpublished rule from a sample of real numbers is both unreliable and a poor use of anyone’s time, since the result is a guess presented with the authority of a specification.
Keep the layers ordered and the messages specific. A user whose input failed on length wants to hear about length, and one whose input passed everything except the closing character wants to hear about the closing character.
For developers: treating country rules as data
The practical consequence of all this variety is that the rules should not be compiled into the code path that decides the verdict.
- Store one record per scheme: the country, the legal character set, the shape, and the level of coverage — algorithm, format only, or none.
- Store the parameters for schemes that publish arithmetic, and leave the field empty rather than guessing for the others.
- Route each input through the records it matches and return a list, not a single winner.
- Keep the four verdicts distinct in your types, so a caller cannot accidentally treat a format-only result as an approval.
- Date the rule set and make updating it a data change, because national conventions change and a hard-coded table becomes wrong silently.
Where a product handles identity values from several countries, the privacy question deserves the same care as the arithmetic. The national ID check digit guide covers how these schemes compare across registers, and CPF CNPJ validation is a worked example of one country whose two tax numbers follow different internal rules.
Next steps
List the countries your product accepts identity numbers from, and mark each one as algorithm, format only or unpublished. Then confirm that your interface says something different for each of those three cases; the number validation tool demonstrates the wording for each level of coverage, and it is worth checking that no message renders a format-only result as an approval.