Number validation is the routine that decides whether a string of characters is allowed to stand for an account, a document or a product. It is not one test but three, applied in a settled order: is every character legal, is the string the right size, and does the closing digit agree with the arithmetic of the characters in front of it.
Each layer catches a different kind of mistake, and each one answers a much narrower question than most people assume. This guide walks through the three layers in order, explains why a passed check is not a claim about the world, and shows where the four possible verdicts of a validator come from.
What is number validation?
At its simplest, validation is a filter that separates input a system can process from input it cannot. The filter is built from published rules rather than secrets, which is why the same check can run in a browser, in a nightly import job and beside a printed label without any of them disagreeing.
The word is often stretched to cover two unrelated activities. One is structural: does this string obey the format its scheme defines? The other is factual: was this number ever issued, and to whom? Only the first belongs to a validator. The second requires a lookup against the register that issued the value, and an offline routine cannot perform it.
Keeping those two apart is the whole point of the exercise. A form that reports a passed structure as an approved identity will wave nonsense through with confidence, and the user will believe it.
Every number and string described in this article is a deliberately constructed illustration, written only to show how the three layers behave. None of them stands for a real account, document or parcel, and a passed check in an illustration is never evidence that any such record exists.
Three layers of filtering: character set, length and check digit
The layers are cheap to run, and each one is blind to the failures the others catch.
| Layer | What it asks | What it catches | What it misses |
|---|---|---|---|
| Character set | Are all characters legal here? | Letters typed into a numeric field, stray symbols, formatting marks pasted in from elsewhere | A wrong digit that is itself legal |
| Length | Is the string the size the scheme allows? | A character dropped or repeated while typing | A same-length string with two neighbours swapped |
| Check digit | Does the closing digit follow from the rest? | Most single-character slips and many adjacent swaps | Any value that was never issued in the first place |
Order matters as much as content. Normalise first, then check the character set, then the length, and only then run the arithmetic. A routine that computes a check digit before removing spaces will refuse values that are perfectly good, and it will do so with a message pointing the user at a problem that does not exist.
Why does passing the format not make a number real?
A check digit is an arithmetic summary of the characters that precede it. It is computed from those characters alone, so it proves internal consistency and nothing beyond it. Anyone who understands the published rule can produce a string that satisfies it, which is precisely what a generator does on purpose.
What the check cannot see is the register. Whether an account is open, whether a document was ever printed, whether a product was ever packed — all of that lives in a database under an authority the validator has no access to. A value can satisfy every offline layer and still correspond to nothing at all; the check digit guide works through what that means for identification schemes in particular.
Treat a verdict as a statement about the string, never about the world. That framing keeps error messages honest and stops reviewers reading more into a green tick than it can carry.
One number can belong to several schemes
Schemes overlap. Two independent numbering systems can agree on the legal character set, the total length and even the arithmetic of the closing character, so the same string is a true member of both. Nothing about that is a defect; it is what happens when separate designers reach for similar conventions.
A validator that announces a single guessed country hides the ambiguity and invents information it does not have. A better interface lists every scheme the string genuinely satisfies and lets the reader decide which register to consult next. The same reasoning runs in reverse when nothing matches: the honest answer is that no implemented rule recognises the input, not that the value is false. A perfectly genuine number can simply fall outside the set of rules a tool knows about, which is the whole subject of numbers without check digits.
Normalisation: spaces, hyphens and case
Real input arrives formatted for human eyes. Account numbers are printed in groups, identity values carry dots and slashes, and labels mix upper and lower case. None of that is part of the number, and all of it has to go before any comparison happens.
Normalisation strips the separators, collapses runs of whitespace and folds letters to one case. Its effect is that two differently formatted copies of a single value become the same string, which is what makes duplicate detection and equality tests meaningful in the first place.
Two cautions are worth keeping. Normalising is not a repair, because it removes presentation rather than mistakes, so it will never rescue a mistyped character. And it is not a security measure either: a routine that compares normalised strings is still comparing untrusted input.
For developers: how to arrange a validation pipeline
Model the pipeline as a sequence of pure steps, each returning a result together with a reason, rather than as one boolean that swallows every distinction.
- Normalise the raw input and keep the normalised form beside the original.
- Verify the character set and refuse anything outside it before any arithmetic runs.
- Verify the length against the scheme the caller selected.
- Compute the check character and compare it.
- Return a verdict that separates a failed check from an unknown scheme.
Keep the reason, not only the outcome. A field that says merely that the value is invalid forces the user to guess, while one that says the format matched but the closing character did not tells them what to change. Where no rule exists for the value at all, say that plainly instead of reusing the failure wording; the check digit algorithms overview shows how much variety sits behind the last layer, and if the field holds a card number the card-specific rules belong to the Luhn algorithm guide rather than here.
Finally, keep the rules as data. When a scheme changes its length or its arithmetic, an update should be a change to a description of the rule, not a change to the logic that reads it.
Next steps
Take one numeric field in your product and write down which of the three layers currently guards it. Then paste a value that satisfies the format but fails the arithmetic into the number validation tool and read the verdict it returns; the difference between a failed check and an unrecognised scheme is the distinction most forms still blur.