Menu

Career profile test data: building CV records

Career profile test data is a synthetic employment record covering roles, education, skills and credentials. Here is what it contains and why coherence matters.

Published

  • test data
  • career
  • software testing

Career profile test data is a made-up working life, written down in full. It is a current role and the company behind it, a work history stretching back a decade or more, an education history attached to it, a set of skills, and the credentials that industry recognises — assembled so that hiring software can be exercised on a record that reads like an ordinary candidate without describing anyone who exists.

This article covers what such a record contains, the situations that genuinely call for one, why a real CV is a poor substitute even when you have permission to use it, and where the internal rules of a career record actually live.

What is career profile test data?

It is the employment-shaped companion to a test identity record. Where an identity record answers the questions a registration form asks, a career profile answers the questions an application form asks: what do you do now, where have you worked, what did you study, what can you do, and what have you been certified in.

The record is not a blob of text and not a single field. It is a small collection of related tables that together describe a working life, and the collection has three properties that make it usable rather than merely present. It is internally consistent, so a reader cannot catch two fields contradicting each other. It is grounded, so the company, the qualification and the credential all belong to the country and the industry the record claims. And it is synthetic, so nothing in it was copied from a person and nothing in it belongs to one.

That last property is the one people underestimate. A grep for obvious placeholder text is not what makes test data safe; provenance is.

When does a team actually need records like these?

Six situations come up repeatedly. They look unrelated, but each one needs plausible input that lands somewhere harmless.

Situation What breaks without generated profiles
Application and onboarding form testing Multi-step forms cannot be walked end to end, because no record carries every step
CV import and parsing work A parser fed the same string twice never meets the section orders real documents use
Staging database seeding Empty tables hide query plans, pagination defects and index mistakes
Product demos and screenshots Screens that read as placeholder text make the product look unfinished
Search, filtering and matching features A recruiter’s saved search cannot be shown matching anything
Access, scoping and export review Nobody can be shown the right fields for a role without a record of that shape

The common thread is that each of these needs input which resembles what the system will eventually receive. A single record asserting nothing in particular is enough for none of them.

Why is copying real CVs the wrong move?

Because a CV is personal data with an unusually long reach. It carries a name, an employment history, an education history, contact details and often a salary expectation — and moving that into a development environment creates a second copy of it in the place where controls are weakest.

Two costs follow. The first lands in whatever privacy regime governs you, since the data was collected for a hiring decision and is now serving a different purpose. The second is quieter and purely operational. The copy spreads: into backups, into query logs, into CSV exports on laptops, into screenshots pasted into tickets, into whatever analytics tool someone pointed at the database for a week. Deleting the source row does not reach any of those, and none of them are audited.

There is also a technical cost. A handful of borrowed CVs can only ever show you a handful of shapes, and real candidate populations are full of shapes a small sample misses — gaps that need explaining, overlapping part-time work, qualifications earned long after the first job, and candidates with no formal qualification at all. Because these records were never anybody’s, they can be generated across that whole spread deliberately.

What does a career profile record contain?

On this site, the career profile generator builds the whole profile at once rather than a field at a time, and the identity key tying it together is the same one the identity page uses.

  1. Current role — the title held now, the company, and the total years of experience that follow from the history below.
  2. Work history — each past role as a title, a company and a start and end month, with the open end carrying the word Present instead of a date.
  3. Education — institution, qualification, field of study and graduation year, with the highest qualification leading.
  4. Skills — the abilities a reader would expect this role to have, tagged rather than scored.
  5. Certifications — the credentials that are genuinely common in this industry and country, each with the issuer that grants it.
  6. Salary — an optional field, and the only one that is usually left out of a test record on purpose.

Two properties are worth understanding before you lean on the output. The first is that some fields are derived rather than independent: the years of experience follow from the dates, the graduation year constrains when the first role can begin, and the current role appears both at the top of the page and as the newest entry in the work history. The second is reproducibility — the same identity key yields the same profile, which is what lets an automated test assert on a specific value instead of merely checking that something arrived.

Why must the record hold together across fields?

Because a career record is mostly a set of relationships, and a defect in a relationship is invisible to a test that only inspects values.

Every field can look perfectly reasonable on its own while the collection is nonsense. The graduation year can be plausible and the first job can start before it. Both roles can carry sensible dates and overlap across the same full-time months. The current role can read correctly and carry an end date, which asserts that somebody who works here has already left. The total experience can be a round number that contradicts the dates directly beneath it.

None of that throws an exception. A parser will happily read a flipped date range into the database, and a form will happily save an education entry that finished after the first payslip. The defect surfaces weeks later, in a report nobody trusts, and the usual diagnosis is that the test data was wrong — which is only half true. The test asserted on values and the data was wrong about relationships, and the test was never written to notice.

A record that is coherent is therefore worth more than a large pile of records that are not. One profile satisfying every rule exercises more of the machinery than a thousand rows that violate it, because uncoherent data skips the interesting code paths instead of testing them.

For developers: designing the record and its dependencies

Model the profile as a small graph of records with dependencies, not as a flat row of independent columns.

Treat a role as a job entry carrying a start date and an end date where the end may be genuinely absent, and derive anything else — the duration, the total experience, the ordering — from those dates rather than storing it alongside them. Store the education entry as a graduation date, and derive the graduation year from it; the moment the year is stored twice, the two copies will disagree. Store a salary as an amount with a currency code and a pay period, never as a bare number. Where a field cannot exist for a given country and industry, leave it absent rather than filling it with a plausible-looking string, because an empty field and a wrong field fail in completely different places.

Then decide the generation order, because the order is the constraint. Country and industry first, since they determine which titles, skills and credentials are plausible at all. Role and seniority next. Education matched to the role rather than drawn at random. Skills and credentials filtered by the same two choices. A flat generator that draws each field independently will produce a record whose parts disagree far more often than intuition suggests.

For fixtures, freeze the profile you assert on and regenerate on the key rather than re-randomising on every run. And design the record so it could never be mistaken for a real person: throwaway company names, fictitious institutions, and a note shipped with the export saying in plain words that the whole set is synthetic, that it exists for software testing, form demos and data seeding, and that it must not be used to impersonate anyone’s employment history, qualifications or credentials.

Next steps

Take the longest form in your product that collects a career and fill it with exactly one generated profile rather than with placeholder text, then submit it and read what comes back. Look specifically for the relationships rather than the values: does the first role begin after the graduation year, do the entries run without overlapping, and does the current role stay open-ended. An export that only looks right field by field is not yet test data. The timeline rules and the parsing fixtures articles take the two hardest parts of that check apart.

Keep reading

Fake Resume & Job Data Generator guides