Checking one number is an interaction. Checking fifty thousand is a batch process with its own economics: the cost per row has to be small, the failure modes have to be legible without a human reading every line, and the output has to be good enough for somebody to act on the next morning.
The temptation is to loop the interactive routine over the file and print the failures. That works for a few hundred rows and collapses afterwards, because it produces undifferentiated noise and hides the two facts a reviewer actually needs: which rows are wrong, and in what way.
How does bulk validation differ from checking one value?
The differences are all about what happens after the arithmetic. An interactive check returns one verdict to one person who can see the value on screen. A batch check returns a verdict for every row of a file that nobody has looked at yet, and the report is the only thing anyone will read.
Scale also changes the shape of the work. Reading a large file into memory at once will fail on the largest inputs, so the pipeline has to stream. Recomputing shared setup for every row wastes most of the running time, so scheme lookups should be resolved once. And the same row may be processed twice if a job restarts, so the operation needs to be safe to repeat.
Finally, the audience is different. A batch report is read by someone deciding what to fix, which means it has to be ordered, categorised and specific about location. The single-value message that says a check digit failed is technically correct and practically useless in a file with forty thousand rows.
Every value referenced in this article is described rather than quoted. A batch run over real data should have such values masked before anything is written to a report, and the examples here are illustrative shapes only.
Clean first, then validate: the pipeline
The order of operations is what makes the rest of the work possible, and it is the same order as in a single-field check, only applied per row.
- Read the file as text, preserving the original values exactly as supplied.
- Normalise each value — strip separators, collapse whitespace, fold case — and keep both forms.
- Resolve the scheme or schemes for each row, based on the column the value came from and, where the column is mixed, on the value itself.
- Run the appropriate check, or record that no rule applies.
- Assign each row to a category, then write the report.
Normalising the whole file before validating means one code path handles both the check and the duplicate analysis. It also means the report can show both forms side by side, which is the fastest way for a reviewer to spot a systematic problem such as a whole column arriving with an extra separator.
Do the scheme resolution once per distinct pattern rather than once per row. A file of one kind of identifier only needs one lookup, and even a mixed file usually contains a handful of distinct shapes, so caching the resolution turns a per-row cost into a per-file cost.
Which categories should a result set contain?
The categories are what turn a list of failures into a diagnosis, and they should match the verdicts the validator already produces rather than invent a new vocabulary.
| Category | What it means | Typical cause |
|---|---|---|
| Checked and valid | The published algorithm was applied and the trailing character agrees | Genuine values, or values generated for testing |
| Checked and invalid | The algorithm was applied and the trailing character disagrees | A typo in the body or in the closing character |
| Format only | The shape was confirmed; no algorithm is published | A scheme in the middle coverage tier |
| Unknown scheme | Nothing the tool implements matches the shape | A mixed column, a stray header, or a value from another system |
| Duplicate | The normalised value appears elsewhere in the file | The same record exported twice |
Duplicates deserve their own category even though they are not validation failures. In an import they are often the single most consequential finding, and burying them among malformed rows guarantees they will be missed.
Keep format-only and unknown separate as well. They lead to different actions: a format-only result means the data is as good as it can be judged, while an unknown scheme needs someone to find out what the column actually contains.
Duplicates, empty cells and very large files
Three ordinary situations account for most of the difficulty in real files.
Duplicates must be detected on the normalised value, because two differently punctuated copies of one number are the same number. Report the first occurrence as the location and list the others as references, so a reviewer can see whether the repetition is accidental or structural.
Empty cells are not failures. A blank where a value was required is a completeness problem, and folding it into the invalid category will inflate the error count and send someone looking for a check-digit bug that does not exist. Count blanks separately, and let the consumer decide whether they are acceptable.
Large files need streaming and a stable memory profile. Read row by row, hold only the deduplication index and the counters, and flush report rows in batches so the output does not become the bottleneck. If the deduplication index itself grows beyond memory, fall back to a sorted external merge or a temporary store rather than trying to hold everything at once.
What makes a report actionable
A reviewer at nine in the morning needs four things from the output: the row number, the value as it arrived, the category, and a short reason. Everything else is decoration.
Row numbers must refer to something the reviewer can find. Say plainly whether they are file lines or data rows, because a header offsets them by one and a reviewer who trusts the wrong convention will edit the wrong record. Include the original value exactly as supplied, since that is what they will search for, and include the normalised form when it differs.
Order matters too. Grouping by category puts every instance of one problem together, which lets a reviewer recognise a pattern instead of reading forty thousand lines. Within a category, order by row so the file can be edited top to bottom.
Finally, make the summary honest. A count of valid rows must say which verdict produced it, because a file where most rows are format-only has been confirmed in shape only, and a summary that reports them as valid will be quoted out of context.
For developers: chunking, concurrency and idempotence
The batching layer is where a working script becomes a dependable job.
- Process in chunks sized by memory rather than by round numbers, and make the chunk size configurable.
- Parallelise the arithmetic, not the output. Each worker should return results, and a single writer should assemble the report so ordering stays deterministic.
- Make the run idempotent: the same input file must produce the same output, including ordering, so two runs can be compared.
- Record the rule set version and the input checksum in the report header, so a result can be reproduced months later.
- Never write full values into logs. Reports are destinations for values, logs are not, and the masking rule should be applied before anything leaves the process.
Masking is not optional when the file contains real values. A bulk run reads everything at once, so a single unredacted report is a much larger exposure than a single failed form submission. Decide before the first run which columns may be reproduced and which must be truncated. The link between a batch run and the API that feeds it is worth reading next: validation at API boundaries covers what the receiving service should do when the file arrives, and how number validation works is the per-row logic this pipeline is built from.
Next steps
Run your current process over a file that deliberately contains one of each category — a valid row, an invalid row, a format-only row, an unknown shape and a duplicate — and check whether the report makes all five obvious without opening the source file. The number validation tool is a convenient place to confirm the wording of each verdict before you fix it in the report format.