Synthetic data vs anonymized data sounds like a choice between two flavours of the same thing, and treating it that way causes real trouble. One of the two was never anybody’s; the other was somebody’s and has been processed to remove them. The difference decides who may see the data, how long it may be kept, and whether it can be shared outside the organisation at all.
This article sets out the four approaches people actually use for test data, what each one does and does not guarantee, and why combining fields is the detail that quietly defeats most attempts at removal.
The four approaches, side by side
Most test datasets are produced by one of four methods, and they are frequently described with the same word.
| Approach | How it is made | What it guarantees |
|---|---|---|
| Synthetic | Values are generated from rules and randomness | Nothing in the set ever described a real person |
| Anonymised | Real records are processed so individuals cannot be identified | The original data existed and the transformation must be proven |
| Pseudonymised | Direct identifiers are replaced by codes, with a mapping kept | The data is still linkable to people through the mapping |
| Masked | Sensitive characters are replaced for display | The underlying value usually still exists |
Only the first starts from nothing. The other three all begin with real records, which is why each of them inherits an obligation that the first one never had.
Why does removing names not anonymise a dataset?
Because a name is only one of many ways to single out a person, and usually not the most reliable one. A table with the name column dropped still tells you that a woman in her thirties works at a particular small employer in a particular small town, and in that population there may be exactly one such person. Nothing has been anonymised; the identifier was simply replaced by a set of identifiers that take slightly more effort.
The statistical version of the same problem is that any dataset carries quasi-identifiers — values that are not unique on their own but become unique in combination. Birth date, postcode, gender and occupation are the classic set, and the arithmetic is unforgiving: a combination that looks broad across a whole country can be unique within a city, and a test database is usually a slice of one market rather than a sample of the world.
Re-identification is therefore not exotic. It is the ordinary consequence of combining a few columns against publicly available information, and it gets easier with every additional dataset that becomes available. That is why an honest anonymisation claim has to address what remains, not merely what was deleted.
What actually counts as anonymous?
Data is anonymous when the individuals in it cannot be identified by anyone, using any means reasonably likely to be used — which is a stronger condition than most teams assume, because it includes information held elsewhere. Practically, that means a proper anonymisation claim rests on a documented assessment: which fields were removed or generalised, what the resulting population looks like, what remains in combination, and why identification is not reasonably possible.
Two consequences are worth taking seriously. First, the assessment is specific to a dataset and a context, so an approach that works for one release may not work for the next if new columns are added. Second, the assessment is genuinely hard, because proving a negative about identification is much more work than demonstrating that a string column was dropped.
Where the difficulty is high, the honest answer is usually to stop calling the data anonymous. Pseudonymised, restricted, or internal-only are all defensible descriptions; anonymous is a claim that has to be earned.
Which approach should a test environment use?
Synthetic data, in the great majority of cases, because it removes the obligation instead of managing it. A dataset generated from rules and a fixed input has no individuals in it at all, so there is no retention clock, no consent question, no removal request to satisfy, and no breach to report if it leaks. It can be committed to a test repository, emailed between teams and pasted into a ticket without a second thought about whose information it contains.
It also sidesteps a practical problem that disguised data keeps generating: the mapping. Pseudonymised data needs a mapping kept somewhere, and that mapping is itself a sensitive dataset that has to be protected, rotated and audited. Teams routinely spend more effort protecting the mapping than they would have spent generating the data.
What synthetic records do not do is automatically reproduce every property of real data. If a test depends on the distribution — the frequency of a rare edge case, the correlation between two fields, the shape of a long tail — that has to be modelled deliberately rather than inherited. The field consistency requirements are a case in point: a generator must be built to keep related fields in agreement, because randomness alone will not.
How do you keep a generated dataset reproducible?
By making it a function of an input you control rather than of the clock or of a random source you cannot replay. The usual arrangement is a seed: a short, human-recordable value that drives every decision the generator makes, so that the same seed produces the same records on any machine and on any day.
Reproducibility matters for three reasons that are easy to forget under time pressure. An automated test can hold a fixed sample and assert against it, so a failure is diagnosable rather than a coin flip. A bug report can name the exact record it saw, so anyone can rebuild it on their own machine. And a review can reproduce a demonstration exactly, which matters when a dataset is being offered as evidence that no real records are involved.
Generated records from the identity and test data generator work this way, and the values follow each country’s real formatting conventions while remaining entirely invented. They are synthetic records intended for software testing only; they do not describe a real person, and they cannot be used to stand in for one or to pass any real verification.
How do you prove a test dataset holds no real records?
By being able to describe where every value came from. That is a provenance question, and it is answered at generation time rather than at audit time. A dataset produced by running a generator with a known seed, from a generator that reads no production source, has a one-line answer. A dataset produced by transforming a production export has an answer that depends on the transformation being correct, and the burden of proving that never quite goes away.
Two habits make the provenance durable. Label the data itself — a column or a header marking rows as synthetic — so that a copy which travels into a ticket, a spreadsheet or a screenshot still announces what it is. And keep the generation step in version control, so that producing the same dataset again is a command rather than an oral history.
There is a final check worth applying to any dataset that claims to be safe: try to identify one person in it. If the attempt succeeds even once, the claim was overstated, and the correct response is a narrower claim rather than a louder one.
For developers: seeds, distributions and honest claims
Design the generator so that every random decision flows from the seed through a single documented path. Two generators that both accept a seed but consume randomness in different orders will disagree, and the disagreement will look like a bug in the test rather than in the generator.
Model the distribution deliberately. Real data has uneven amounts of everything, and a uniform draw produces a dataset that is too tidy: no rare names, no unusual ages, no absent fields. If the test needs a long tail, the generator has to produce one on purpose, and that is a specification question about the data rather than a property that emerges.
Keep the claim about the data narrow and true. Synthetic and generated for testing is a claim that can be verified by reading the generation step. That is worth far more to an auditor than a broader claim about anonymity that nobody can reproduce. And carry the constraint into the fixture notes: records produced this way exist to exercise software, not to impersonate anyone or to be offered anywhere as a real identity.
Next steps
Find one dataset in your test environments whose origin nobody can state, and trace it; that answer is usually more interesting than the dataset itself. Then replace it with a generated batch and record the seed alongside the data, so that anyone can rebuild it. If you are weighing the alternatives, the privacy rules for test data article covers what each choice obliges you to do afterwards.