HR data retention is the discipline of deciding how long each kind of personnel record should exist, and then making the system actually do it. It is difficult for a reason that has nothing to do with technology: the answer is different in every jurisdiction and for every category of record, and no single figure can be correct everywhere.
This article sets out why there is no universal period, how personnel records split into groups with different clocks, which principles govern the decision, and what deletion has to cover to mean anything.
Why is there no single retention period?
Because the period is set by whoever has authority over the record, and that varies by jurisdiction and often by the type of record within one jurisdiction.
Some periods exist because a claim can still be brought for that long, and the period is set so that the evidence survives the claim. Some exist because a regulator requires the record for inspection. Some exist to serve the person the record is about, so that they can obtain a reference or prove a period of employment years later. And some exist merely because the system has never been told to stop keeping them, which is not a rule at all but the absence of one.
The practical consequence is that anyone designing retention should treat the period as configuration supplied by the operator for their own jurisdiction, not as a constant to be chosen by the developer. A hard-coded duration is wrong in most of the places it will run, and it is wrong silently, because nothing in the system checks whether the number still applies. Where a period is configurable, the operator can answer for it; where it is baked in, nobody can.
How do personnel records divide by purpose?
By what the record is for, because that is what determines its clock. Four groups cover most of what a hiring system holds.
| Group | What it is for | Why its clock differs |
|---|---|---|
| Active employee records | Administering the employment | Kept while the relationship lasts, and usually beyond it |
| Former employee records | References, claims and statutory duties | Kept for a period after the relationship ends, set by the operator |
| Unsuccessful candidate material | Evidence about a decision | Usually the shortest-lived, because the decision is closed |
| Interview notes and assessments | The reasoning behind a decision | Frequently overlooked, and the group most often kept informally |
The last two groups cause the most trouble, and for the same reason: they are created in the middle of a process, by people who are not thinking about a records regime, and they often live outside the system that manages everything else. Notes taken in a meeting, messages between interviewers and a spreadsheet used to compare candidates are all personnel records. A retention policy that covers only the applicant tracking system is covering part of the picture.
Treating all four groups as one record with one expiry date is the most common design error. It produces the worst of both outcomes: candidate material kept far longer than the decision requires, and employment records expired earlier than the operator’s own obligations allow.
Which principles should govern the decision?
Six, and they are the ordinary ones for personal information, applied to a hiring context.
- Purpose limitation — collect the record for a stated purpose and do not repurpose it later because it happens to be there.
- Minimisation — hold what the purpose requires and nothing that merely might be useful.
- Accuracy — keep the record correct, and give the person a route to correct it.
- Storage limitation — define how long each record lives, and enforce it rather than assuming it.
- Integrity and confidentiality — limit who can reach the record, and know who did.
- Accountability — be able to show that the first five are actually happening.
Two of these are frequently asserted and rarely implemented. Storage limitation tends to exist as a written policy with no mechanism behind it, so records outlive their period and nobody finds out. Accountability tends to mean a document rather than a practice, so there is no way to answer a question about what was deleted and when.
Minimisation deserves a specific warning in a hiring context, because the temptation runs the other way. Extra information feels harmless at collection time and is exactly the material that creates exposure later. A field that is not needed for the decision is not a neutral addition; it is a liability with a retention clock attached.
What does deletion have to cover?
Everything the record touched, which is almost never a single table.
A record’s life leaves traces in the places it passed through, and each of those places needs a rule. Live storage, backups and archives, exports produced for reporting, log entries that contain field values, search indexes built from the record, cached copies held by an integration, and any copy a person downloaded to a local device. Deleting the row and stopping there is a common and serious undercount.
Three distinctions make the design tractable. Separate deletion from anonymisation, and be honest about which one is being done — removing the link to a person is not the same as removing the record, and an anonymised record that can be re-linked has not been anonymised. Separate expiry from request-driven deletion, since one runs on a schedule and the other arrives at an arbitrary moment and must propagate through the same places. And separate the record from the evidence that the deletion happened, which normally needs to survive the deletion itself.
Backups are the part that resists all of this, because the ordinary mechanism for recovering the latest state is the same mechanism that preserves the deleted record. The usual reconciled position is that backups are excluded from immediate deletion but must be covered by an expiry that eventually reaches them, and that a restored backup re-applies the deletions made since it was taken. Whichever position is taken, it should be a decision that is written down rather than an omission nobody noticed.
For developers: retention as configuration and purge
Treat the period as data. Give each category of record its own configured period, supplied by the operator, and let the system enforce whatever it is told rather than carrying an opinion.
Five habits make the difference. Tag every record with the category that determines its clock, so a purge can find records by rule rather than by a hand-maintained list. Attach the clock to the category rather than to the table, because the same table holds records of several kinds. Log the purge itself — when it ran, what it matched, and what it deleted — and keep that log where the purge cannot remove it. Make the fields you collect the ones your documented purpose needs, so minimisation is a property of the schema rather than a promise. And give every environment that holds personnel data, including test and demonstration environments, the same regime, because a copy in a test environment is a copy.
That last point is the one to act on first. Personnel data should not be used to populate test environments at all. Sample records should be constructed for the purpose, and where a copy of real data is unavoidable for a specific investigation, it should be governed as though it were production. Everything on this site is generated for testing and demonstration, and the personnel records it produces are intended to replace real ones in exactly those environments, never to supplement them.
Next steps
Find out whether your systems can answer one question: which candidate records from two years ago no longer have a purpose, and what is still holding them. The answer is usually a list of places nobody thought of, and it is a better starting point than a policy document. How a synthetic career record is assembled in the first place is covered in career profile test data, and one of the most sensitive fields in such a record is treated separately in salary currency and period. The career profile tool produces records intended for test and demonstration use rather than for production.