Menu

Resume parsing fixtures that stay useful

Resume parsing fixtures need to cover layout variants rather than file numbers. Here is how to organise them, what they should assert, and how they decay.

Published

  • test data
  • career
  • parsing

Resume parsing fixtures are the small library of documents a team keeps so that a parser can be tested against something other than the file somebody happened to have open. They are the least glamorous part of a hiring system and the part that decides whether parsing defects get caught in a test run or in production.

This article covers why the format problem cannot be solved by picking a standard, which layouts actually break parsers, how to organise the fixtures so a failure points at a cause, and why a set of fixtures is never finished.

Why is there no standard format for a CV?

Because nothing in the process requires one. A CV is a document a person writes to be read by a human, in whatever tool they have, and the tools impose no shared structure at all. A word processor file with a two-column layout, a text export from an online profile, a scanned printout, a slide-shaped deck — all of these arrive at the same inbox and all of them are legitimate.

The absence of a standard is not an oversight that better tooling will close, because the incentives run the other way. The person writing the CV is optimising for the impression a human reader forms in the first few seconds. The person reading it is optimising for whatever they can extract quickly. Neither of them has any reason to constrain the layout to suit a parser.

That is the premise the whole testing strategy rests on. The goal is not format compliance, because there is no format to comply with. The goal is recalling the right fields from documents whose layout nobody controls.

Which layout variants should parsing tests cover?

The ones that move content around. Order, columns and proximity between a label and its value are what break extraction, far more often than unusual fonts or a different page size.

Layout variant What it breaks
Two-column body Reading order is lost, so a right-column date attaches to a left-column role
Section titles in an unusual order A parser that expects education first records nothing for the entries that come later
Experience written as prose No repeated block structure, so a parser that needs one finds none
Skills as a tag row Items run together without separators, and the split between them is invented
Tables used for layout A role, an employer and a date sit in three cells with no semantic relationship
Dates in a non-numeric form Month names, seasonal labels and approximate wording defeat numeric matching
Headers and footers Contact details are duplicated into every page and can overwrite the real ones

A second axis matters just as much: which pages the contact details live on, and whether the document has one variant or two. A CV exported twice from the same tool with a language switched will have the section titles in a different language, and a parser that keys on those titles will silently extract nothing from the second export.

Should fixtures be organised by scenario or by number?

By scenario, always. The number tells you nothing when a test fails, and the scenario tells you almost everything.

A fixture named for what it contains — a two-column layout, a tag row of skills, an experience section written as a paragraph — is self-documenting. When it fails, the name is a hypothesis about where the parser broke. A fixture named with an index or a date is a lookup that has to be resolved through a second document before any debugging can start, and that second document is always out of date.

The organisation also determines what is missing. A library sorted by scenario shows its own gaps at a glance: if there is no fixture for a document with no education section, then that case has never been tested, and the absence is visible. A library sorted by file number hides the same absence completely.

There is a second benefit that matters for maintenance. Scenario names survive changes to the parser. When the implementation is rewritten, the fixture library still describes the shapes real documents take, which is the part that does not change.

Why does random generation make failures unreproducible?

Because a random input is not a record of anything. When a test fails against a randomly generated document, the failure exists in one run and nowhere else, and the only way to investigate is to modify the generator so that it reproduces the case — which is to say, to turn it into a fixture.

That is not an argument against random generation, which is genuinely good at finding unexpected shapes. It is an argument about which side of the boundary each tool belongs on. Random generation belongs in the exploratory phase, where the goal is to discover a case nobody thought of. The moment a case is discovered, it stops being random and becomes a fixture, with the input frozen and the expected output recorded. The regression is then locked to a case that exists.

There is a subtler problem with advancing random generation into the regression suite. A randomly generated document varies in every dimension at once, so a failure arising from it cannot be attributed to any single one. A scenario fixture varies one dimension deliberately, and the failure it produces names its own cause.

How do fixtures decay?

Slowly, and in three ways that are easy to miss because each one looks like nothing at all.

The real documents they imitate change: templates get redesigned, a layout that was common five years ago stops appearing, and a new one takes its place. The fixture keeps passing and keeps covering a shape nobody submits any more. The languages drift: a fixture set built in one language tests label matching in that language only, and adding a second language is not a change to a fixture but the addition of a whole parallel set. And the expectations rot against the parser: an assertion written against an early version of the extraction logic may encode a behaviour that was later corrected, so the test now enforces a known bug.

None of these are detected by the suite itself, because every fixture still passes. Decay is found by review: reopening the library periodically and asking whether each fixture still resembles something real, whether every supported language has one, and whether each assertion still describes the behaviour you want.

For developers: fixture shape and expected fields

Keep the input and the expectation together, and keep the expectation narrow.

A fixture is a pair: the document, and the fields the parser should return from it. Keeping them in the same place means a change to an expectation is reviewed next to the shape that caused it. Keep the expected value readable rather than encoded, so a reviewer can see that a date is expected to land in a particular month without resolving a serial number first.

Four practices prevent most of the pain. Assert on the fields the layout is designed to stress, and leave the rest of the extraction loosely checked, so a fixture is not broken by an unrelated improvement elsewhere. Record the origin of every fixture in a line of prose — constructed for this purpose, layout modelled on a common shape, no real document used — so that nobody later has to guess whether a file came from somewhere sensitive. Version the fixtures with the parser so it is possible to tell which set a given result was produced with. And keep a deliberate gap in the library for the case you know is unsupported, documented rather than silently absent, because a known gap is a decision and an unknown one is a defect.

Everything a fixture holds should be constructed for the purpose. The sample documents and expected fields described here are built samples that contain no content taken from any real CV, and they exist to exercise parsing logic rather than to represent anybody.

Next steps

Choose three fixtures from your existing set and rename them so that the name describes the layout rather than an index. The exercise almost always reveals that two of them are the same scenario twice and that an obvious one is missing entirely. The hiring form cases that consume the parsed output are the other half of the same test surface, and if the goal is volume rather than layouts, the walkthrough on seeding a staging database covers it at scale. When you need documents to build the library from, the career profile tool produces the underlying record whose fields the fixtures should expect.

Keep reading

Fake Resume & Job Data Generator guides