Menu

GDPR Test Data: Why Real People's Details Do Not Belong in Tests

GDPR test data questions start with one habit: keeping real personal data out of development and staging. Here is why, and how synthetic records replace it.

Published

  • test data
  • identity
  • privacy

GDPR test data is a phrase that usually arrives with an uncomfortable realisation attached: the staging database everybody has been querying freely contains real names, real addresses and real identification numbers, and nobody ever decided that should be allowed. It was just convenient.

This article covers what counts as personal data in a test record, why development environments are the wrong place for production copies, the difference between disguising data and removing it, and what a workable replacement looks like.

What counts as personal data here?

More than people expect, and the list is behavioural rather than technical. A name, an identification number, a date of birth, a home address, a telephone number and an email address are the obvious members. Less obvious ones join them: a photograph, a device identifier, a location trace, an account nickname, and any note in a support system that happens to describe someone.

  • Obvious members: a name, an identification number, a date of birth, a home address, a telephone number and an email address
  • Less obvious ones: a photograph, a device identifier, a location trace, an account nickname
  • Also personal: any note in a support system that happens to describe someone

The concept is defined by reference to the person rather than to the field. Data is personal when it relates to an identifiable individual, directly or indirectly, which means a field does not have to contain a name to be personal. A record with a birth date, a postal code and a job title can be personal if those three together pick out one person in the population the data covers.

That definition is what makes the test environment a live question rather than a formality. If the staging database holds any of this, staging is processing personal data, and every argument about retention, access and security now applies to a system that was designed without them.

Why shouldn’t test systems hold production data?

Because the controls that protect production are precisely the controls that lower environments lack. Production usually has restricted access, encrypted storage, audit logging and a change process. Staging typically has none of the four, because it exists to be convenient. Copying a dataset into it is therefore a downgrade in protection applied to the most sensitive records the organisation holds.

Three consequences follow, and each one is expensive in its own way. Access multiplies: every contractor, every automated test account and every laptop that restores a dump now holds personal data. Retention becomes accidental: the copy is refreshed, forgotten, backed up and left in a bucket nobody owns. And incident scope grows: when a routine leak happens somewhere, the question is no longer which systems were exposed but how many copies exist.

The regulatory angle is narrower than the engineering one but points the same way. Personal data must be collected for specified purposes and not used for incompatible ones, so quietly reusing production records to test a migration is a new purpose nobody agreed to. None of this is unique to one legal framework; the same reasoning appears in most privacy regimes, which is why the practical advice converges.

What is the difference between anonymisation and pseudonymisation?

These two are constantly treated as synonyms and they are not. Pseudonymised data has had direct identifiers swapped for a code, with the mapping retained somewhere. The record cannot be attributed without that extra information, but the link still exists and can be followed if the mapping leaks. Anonymised data has been processed so that the individual can no longer be identified by anyone, including the organisation holding it, and the transformation cannot be reversed.

The distinction decides whether the data is still personal data. Pseudonymised records remain personal for most purposes, because a key exists. Anonymised records, done properly, do not — but proving proper anonymisation is genuinely hard, and the difficulty is what makes using production data for testing such a poor trade.

There is a second trap beyond the direct identifiers. Removing names from a table does not anonymise it, because combinations of remaining fields identify people. A rare job title in a small town is often enough on its own. Under a legal test, the right question is not whether a name is present but whether anyone could reasonably identify the individual from what remains, using information they have or could obtain.

Do masking and deletion solve the problem?

Masking replaces sensitive characters with a fixed pattern, and it is excellent for display: a support agent seeing the last four digits of a number while the rest is hidden. It is not anonymisation, because the original value is usually still stored, and a masking layer in front of a real record does nothing about the real record.

Deletion of individual fields has the same weakness in a different costume. Take the name out and the record still has an address, a birth date and a telephone number. Replace the address too and the remaining fields may still single out one person in the population. Each removal narrows the risk without closing it, and there is rarely a clear point at which someone can prove the risk has reached zero.

There is also a documentation problem. In practice, teams cannot demonstrate which records were anonymised properly and which are merely disguised, because the transformation happened ad hoc and left no trace. A dataset whose provenance nobody can reconstruct cannot be defended later.

What should replace production copies?

Records that were never anybody’s. This is the argument for generating data rather than harvesting it, and it is why synthetic records solve both the privacy question and the quality question at once. There is nothing to protect, nothing to retain, nothing to leak, and no purpose limitation to argue about, because the data never belonged to a person in the first place.

Generated records also behave better in tests. Values produced by the identity and test data generator are internally consistent — the address, postcode and telephone prefix belong to the same country, and identifier formats follow that country’s conventions — which means a test dataset does not need hand-repairing. They are synthetic records for software testing only, and they are not usable to impersonate a real person or to satisfy any real identity, age or eligibility check.

The honest caveat is that synthetic data does not automatically reproduce every real-world correlation. If a test genuinely depends on the statistical shape of real data, that shape has to be modelled deliberately, and the comparison of synthetic approaches sets out what modelling entails and where it falls short.

How do you keep personal data out of logs and screenshots?

The leak is rarely the database. It is the query log that captured a full record, the error report that embedded the request body, the screenshot attached to a bug ticket, the spreadsheet exported for analysis and left in a shared drive, and the local dump someone took to reproduce a defect on a plane.

Three controls reduce most of it. First, keep real records out of environments where logs are verbose, which is a direct argument for using synthetic data in staging. Second, treat a screenshot or an export as containing whatever the screen contained, and require that test evidence comes from environments holding no real records. Third, make retention explicit: a copy that exists for a purpose should have an end date, and a copy with no end date is not a copy, it is a second production database.

Access review belongs here as well. Knowing who can reach the dataset is only useful if the answer has been checked recently, and the moment a copy is widely shared is when that review stops being meaningful.

For developers: isolation, provenance and proof

Separate the environments rather than filtering one stream of data into the other. A development environment should be populated by a generation step, never by a restore from production, and the connection from development to production should not exist at all — not merely be unused.

Label the provenance of the data in the dataset itself. A marker that says these rows are synthetic and for testing only survives the export, the screenshot and the support ticket, and it tells a future reader everything they need to know about what they are looking at. It is the cheapest safeguard available and the one most often skipped.

Finally, prepare the answer to the question an auditor will ask: how do you know this dataset contains no real records? A generation step that runs from a seed and never reads a production source is a demonstrable answer. A pipeline that transformed a production export is not, no matter how thorough the transformation looks in review. Verifiable absence beats provable disguise every time.

Next steps

Find out whether your staging database was restored from a production dump, and if it was, make the case for replacing it with generated rows before the next quarterly access review. Then walk through one bug report from last month and count how many places a real person’s details travelled inside it — logs, attachments, screenshots — because that count is what the alternative is competing against. When you do move, generate the first batch in the identity generator and label every row as synthetic before it is imported.

Keep reading

Identity & Test Data Generator guides