Brazilian tax numbers come in two shapes, and telling them apart is the first step of any validation. The CPF identifies a natural person; the CNPJ identifies a legal entity. Both close with two check digits rather than one, which lets a moderately more ambitious arithmetic check catch a wider class of typing slips.
The pair is a good case study because the numbers are short enough to check by hand, old enough that their rules are public everywhere, and different enough from each other that a single careless routine will get at least one of them wrong. Here is how they are built and what a validator can honestly say about them.
What are CPF and CNPJ?
The CPF is the individual taxpayer registration, the number a Brazilian citizen is asked for at a bank counter, a hotel desk or an online checkout. The CNPJ is the corporate taxpayer registration, the number printed on every invoice and business card a company issues. One belongs to a person, the other to an organisation, and the registers behind them are separate.
Because they play parallel roles, the two strings are often discussed together and implemented together, and the shared vocabulary hides the fact that their internal layouts are not the same. A developer who reads about one and assumes the other works identically will produce a routine that rejects valid values about half the time.
The string examples in this article are synthetic illustrations built only to show how the two numbers behave. No real registration belonging to any person or company is reproduced or adapted here, and passing an illustration through the check says nothing about anyone.
How the two numbers differ in shape
The obvious difference is that the legal-entity number is the longer of the two, which is a consequence of the larger register and its internal subdivisions. The person number is the more compact form, sized for a single individual with no branches.
The structural difference matters more than the length. The legal-entity format is assembled from parts that identify the company’s registration at more than one level, so its internal fields are not one homogeneous block of digits. The person format is essentially flat: a serial body, a region marker and the two check digits.
Neither format should be described by reciting its exact positions in prose. The published rules are the source of truth, and a table copied into a wiki page ages badly. Treat the layout as data that the code reads, not as knowledge distributed through comments in three files.
The check digit idea: weighted sum, then modulus
Both formats derive their closing digits with the same family of arithmetic, applied twice. Each of the last two positions is computed from the digits before it: multiply each digit by a weight that depends on its position, add the products, divide by a modulus, and convert the remainder into a check digit through a published mapping.
The second digit is computed after the first, and its calculation includes the first digit among its inputs. That chaining is the reason the two cannot be verified independently: changing the second digit’s computation without also feeding in the first produces values that look right in isolation and fail in combination.
The family is the modulus-based one surveyed in check digit algorithms, and the mapping step is where naive implementations drift. When the division leaves a remainder that does not translate directly into a single digit, the published rule says what to do with it, and a routine that omits that step will disagree with the standard on a small but reproducible subset of inputs.
Why are all-same-digit sequences excluded?
Repeated digits are the classic counterexample to any modular check. A string composed of one character repeated defeats simple arithmetic because the weights and the repeated values interact in a way that leaves the total unchanged, so the check digit comes out matching. The published rules therefore exclude such strings outright rather than relying on the arithmetic to notice them.
That exclusion is not a limitation of the algorithm; it is a deliberate statement that a registration number of that shape will never be issued. A validator that implements the arithmetic but skips the exclusion will accept a value that the register itself would never have generated, which is exactly the kind of gap that lets placeholder text sail through a form.
The same reasoning is worth generalising. Whenever a scheme defines a carve-out, the carve-out is part of the rule, not an optional extra, and the place to implement it is beside the arithmetic rather than in the caller.
Format checks and registry status are different things
A passed check means the two closing digits agree with the body. Nothing more. Whether the number was ever issued, whether it is still active, whether it belongs to the person typing it — all of that lives in the tax authority’s register, which an offline routine cannot see.
The register is also the wrong thing to consult casually. Querying someone’s tax status is a different activity from checking that a string is well formed, and it carries legal and privacy consequences that a form validation never does. Building the arithmetic locally and treating the register as a separate, deliberate integration keeps those two concerns from leaking into each other.
Error wording has to carry that distinction. A user who is told their number is invalid may conclude that their registration has a problem, when in fact only the string failed the arithmetic. The message should say that the digits do not agree, and leave the state of the registration out of it.
For developers: cleaning, masking and error messages
A handful of habits keep this pair of schemes out of trouble.
- Normalise the same way for both: strip the presentation marks, keep the digits, and remember that cleaning is not repair.
- Decide the scheme from context, not from a guess. Ask the caller which kind of taxpayer the field expects instead of inferring it from length alone.
- Report which of the two digits failed, or whether both did, because the position of the failure points at where the typo is.
- Never log a full value in production telemetry. Record the verdict and a masked form.
The same care applies to the people reading the numbers rather than the code. A masked form keeps enough leading characters to be recognisable to its owner and drops enough trailing ones to be useless to everyone else, and the national ID check digit discussion of comparison across countries is a useful companion when a product handles several jurisdictions at once. If your product also accepts cross-border account identifiers, IBAN structure shows how a second, independent check rides inside a single string.
Next steps
Write down which of the two schemes each field in your product is meant to hold, then test your routine with a value whose body is legal and whose closing digits are not. The number validation tool will show you how a validator should describe that outcome, and national ID validation rules puts the pair in the wider context of numbering systems that were designed without one another in mind.