Every retail barcode ends with a digit that is not part of the article’s identity at all. It is arithmetic: a weighted sum of the digits in front of it, reduced to a single character and printed so that a scanner, or a person reading the printed numbers under the bars, can tell whether the code survived intact.
EAN and UPC are the two families most people meet first, and they are close enough relatives that the same routine often handles both. The differences are small, positional and easy to get wrong, which is why this article concentrates on where the arithmetic starts rather than on the arithmetic itself.
Where does the last digit of a barcode come from?
The pattern is the familiar one. Each of the preceding digits is multiplied by a weight, the products are added, and the total is turned into a single trailing character. For the barcode families described here the weights alternate between one and a larger value, and the closing digit is derived by taking the sum up to the next multiple of ten.
The purpose is narrow and specific. A barcode is read by a machine that can misread a bar, and the trailing digit is there to make most misreads obvious before the wrong product is billed. It is an error-detection device for a noisy channel, not a certificate of anything.
Because the sum only ever covers the digits in the code itself, the trailing digit can be recomputed from any candidate string. That is what makes the check cheap enough to run on every scan, and it is also why a passing code is no evidence about the product behind it.
The codes in this article are described rather than printed. No real product barcode is reproduced, and a described scheme is a public convention rather than a statement about any item on any shelf.
EAN-13 and UPC-A: two common lengths, different weight starts
The two families share their arithmetic and differ in the details that surround it.
| Property | EAN-13 | UPC-A |
|---|---|---|
| Total length | Longer of the two | Shorter of the two |
| Weight pattern | Alternating, starting from one side | Alternating, starting from the other |
| Where it is seen | Retail items internationally | Retail items in North America |
| Relationship | The shorter code can be carried inside the longer one with a leading digit added | A subset of the same arithmetic family |
The relationship between them is why libraries often expose a single function that decides the length first and picks the weight start second. A routine that assumes one weight pattern regardless of length will compute a check digit that is wrong for every code of the other size — and it will do so confidently, producing a value that only fails when it meets real data.
Names are worth using precisely here. EAN-13 and UPC-A are standard designations for two specific formats, and calling every retail number by one of them is a small habit that leads to a real bug the first time a routine has to accept both.
The idea of weighting from the right
The weights alternate, and the question that decides correctness is which position gets which weight. In these schemes the alternation is anchored at the right-hand end, so the digit immediately before the check character always carries the same weight and the pattern grows leftwards from there.
Anchoring from the right rather than the left is what makes a single routine work across lengths. If the pattern were anchored at the left, adding a digit would flip the parity of every position and change the meaning of the code entirely. Anchoring at the right means the check character’s own neighbourhood is stable and the extra digit at the far left is simply one more term in the sum.
This is the single most common source of off-by-one defects in barcode code. It is also cheap to test: generate a valid code, verify it with the routine, then extend it by one leading digit and verify again. A routine anchored on the wrong side will disagree on one of the two lengths.
What a barcode prefix does and does not say
The leading digits of a retail code are allocated to numbering organisations, and the allocation is published. It is tempting to read them as a statement about where an item was made, and that reading is wrong. Allocation identifies the organisation that registered the number, and a registered organisation can have goods manufactured anywhere in the world.
Physical label printing adds a second layer of looseness. The same article can carry a code printed in one country and be packed in another, and a company may hold registrations under more than one allocation. Treating the prefix as a production-origin marker will produce confident, incorrect sourcing claims.
For a developer this is a warning about scope creep. The barcode can tell you whether the digits are internally consistent, and it can tell you which organisation registered the number if you consult the published allocation. Everything beyond that belongs to a different system with different data.
Does a passing check mean the barcode will scan?
No, and the two failures are worth separating. The check digit guards against the digits being wrong; it says nothing about the printing. Bars that are smudged, truncated, printed at the wrong magnification or on a curved surface can defeat a scanner even though the digits underneath are perfectly consistent.
Nor does a correct code prove the item exists in your catalogue. A well-formed code can refer to an article you have never stocked, or to a code that was never assigned at all. Catalogue membership and arithmetic consistency are independent properties, which is why a received-goods process needs a lookup against its own list as well as a validation step.
The reverse case is just as important. A misread that a scanner corrects silently, or a code that fails the check because the label was damaged, is not evidence of fraud. It is evidence about the label.
For developers: bulk scanning and locating errors
Barcode validation usually arrives in volume, which turns two small design decisions into large ones.
- Decide the length from the digits alone, then select the weight start from the length. Do not rely on a caller to declare which family it is sending.
- Reject illegal characters before the sum runs, so a pasted description does not become a mysterious checksum failure.
- On failure, say whether the length was wrong or the trailing digit disagreed; those two messages send a user to completely different remedies.
- When checking a file, report the row and the value as received, so a warehouse operator can find the label in question.
Each scheme here belongs to the same modulus-based family as the others described in the check digit algorithms overview, which is useful when one routine has to serve several kinds of code. Book and serial numbers follow the same principle with their own quirks, covered in ISBN ISSN check digits.
Next steps
Pick one valid code from each of the two lengths and confirm your routine accepts both, then change a single digit in each and confirm it rejects both. Afterwards, check what your system does when the arithmetic passes but the code is absent from your catalogue — the number validation tool shows how a validator should phrase a result that is about the digits only.